Vinr user audio and video perception evaluation method
By using the ViNR user audio and video perception evaluation method, which combines video, audio, text similarity and network quality evaluation, a comprehensive user perception evaluation is formed. This solves the problem of not being able to accurately locate network problems in existing technologies and achieves a network quality assessment that is closer to user perception.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NANJING HOWSO TECH
- Filing Date
- 2025-07-03
- Publication Date
- 2026-04-23
AI Technical Summary
Existing methods for assessing voice and video quality cannot directly represent users' subjective experience, especially in ViNR audio and video calls, and cannot accurately pinpoint network problems.
The ViNR user audio and video perception evaluation method is adopted. By processing audio and video data at the sending and receiving ends, and combining video similarity, audio similarity, text similarity and network quality evaluation, a comprehensive user perception evaluation is formed. This method closely resembles the human brain's thinking mode and uses time and location mapping to locate network problems.
It more accurately locates network problems, solves the problems of poor repeatability of the MOS subjective evaluation method and the inability of the MOS-LQO objective problem to reproduce the human brain's thinking paradigm, and gets closer to the user's perception of the quality of network audio and video calls.
Smart Images

Figure CN2025106830_23042026_PF_FP_ABST
Abstract
Description
ViNR User Audio and Video Perception Evaluation Method
[0001] This application claims priority to Chinese Patent Application No. 202411431500.0, filed on October 14, 2024, entitled "ViNR User Audio and Video Perception Evaluation Method", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates to the field of communications, specifically to video call services in the field of communications, and particularly to a ViNR user audio and video perception evaluation method. Background Technology
[0003] With the development of wireless communication, in the current 5G era, operators have launched the "5G New Voice" service, which is a high-definition audio and video call service based on the 5G network, using VoNR (Voice over New Radio) / ViNR (Video over New Radio). ViNR audio and video services experience various distortions during acquisition, compression, transmission, and storage; any distortion can lead to a decrease in the user's perceived voice and visual quality.
[0004] Currently, audio and video quality assessment methods in the industry are divided into two main categories: voice quality assessment and image and video quality assessment.
[0005] Voice quality assessment is divided into two types: subjective assessment and objective assessment. ITU-T P.800 defines the subjective testing method for MOS (Voice Quality Assurance), while objective testing methods mainly include PESQ and POLQA. Among them, ITU-T P.863 (POLQA) is currently the method recommended by the ITU for VoLTE voice quality testing.
[0006] Image and video evaluation is divided into two types: subjective evaluation and objective evaluation. Subjective quality evaluation mainly relies on human visual observation and scoring, which can be said to be the most intuitive way to reflect the audience's perception of video quality and is also the ultimate goal of other objective evaluation methods. However, this method has problems such as being time-consuming, labor-intensive, costly, and subject to subjective bias, making it unsuitable for direct application in industry. Objective quality evaluation of image and video typically uses Quality Assessment (QA) algorithms for modeling. QA algorithms can accurately measure the merits of encoding / decoding models, communication transmission systems, image enhancement, and reconstruction algorithms. According to the International Telecommunication Union (ITU) recommendations, they can be divided into five categories based on the data type of input: Media-layer models, Parametric packet-layer models, Parametric planning models, Bitstream-layer models, and Hybrid models. Among them, the Media-layer model directly uses media information for computational analysis to give the evaluation result, while other types of evaluation methods evaluate quality based on external variables such as coding parameters or network channel conditions.
[0007] Chinese patent document CN107920362 A discloses a micro-region-based LTE network performance evaluation method, including the following steps: (1) Data collection: collecting user-level OTT information, MR data, key signaling handover data, and call statistics data; (2) Establishing a location fingerprint database; (3) Data processing: integrating and associating various data sources; simultaneously, classifying the call statistics data under the two major service types of LTE and VoLTE according to five dimensions: maintainability, accessibility, integrity, cell integrity rate, and mobility, and marking the indicator attributes; (4) Data calculation and analysis; (5) Data analysis results: the service type is divided into two types: LTE (browsing service) and VoLTE service. The evaluation time can be selected by the user. The network performance score of the grid is divided into five intervals: excellent, good, average, poor, and severe. By utilizing the correlation and constraint relationships between indicator sets within each dimension, the network quality of the micro-region can be evaluated reasonably and objectively, effectively guiding network optimization.
[0008] Current quality assessment methods evaluate speech quality and image / video quality separately. The numerical values for these individual metrics can vary significantly and cannot directly represent the user's subjective experience, especially regarding audio and video quality. Therefore, it is necessary to develop a ViNR user audio and video perception evaluation method that can more accurately pinpoint network problems. Summary of the Invention
[0009] The technical problem to be solved by this invention is to propose a ViNR user audio and video perception evaluation method, which not only solves the problem of poor repeatability of the subjective evaluation method of MOS, but also solves the problem that the objective problem of MOS-LQO cannot reproduce the human brain's thinking paradigm. It is closer to the human brain's thinking mode and closer to the user's perception of the quality of network audio and video calls. At the same time, through time and location mapping, combined with network parameters and events, network problems can be located more accurately.
[0010] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: the ViNR user audio and video perception evaluation method, which specifically includes the following steps:
[0011] S1 Audio and Video Data Initiation: The sending end initiates audio and video data, records and saves the time-series data corresponding to the audio and video data at the time of initiation; the audio data is processed and saved, and the network data of the sending end is also saved.
[0012] S2 Audio and Video Data Reception: The receiving end receives audio and video data, processes and saves the time sequence corresponding to the received audio and video data; in particular, the audio data is processed and saved, and the network data of the receiving end is also saved.
[0013] S3 Evaluation and Analysis: Similarity evaluation methods are used to evaluate various data from both the sending and receiving ends. Time series data is combined for time series evaluation, and network quality is combined for network data evaluation. Finally, a comprehensive audio-visual perception evaluation is obtained to form a user perception evaluation.
[0014] The above technical solution utilizes video similarity, audio similarity, text similarity, network quality assessment, and time series evaluation for comprehensive audio and video perception evaluation. This solves both the poor repeatability problem of the MOS subjective evaluation method and the inability of the MOS-LQO objective problem to replicate the human brain's thinking paradigm. It is closer to the human brain's thinking pattern and more closely reflects the user's perception of network audio and video call quality. Furthermore, through time and location mapping, combined with network parameters and events, network problems can be located more accurately. The receiving end and the sending end are different clients or platforms. This method processes audio and video data from both the sending and receiving ends, converting the audio data into text information and evaluating similarity using a similarity fitting algorithm for video, audio, and text data. It also evaluates the time series data from the transmitted and received ends using a time series alignment algorithm and method. Furthermore, it displays and saves network parameters and event information of the communication units connecting to the network in real time, and evaluates network quality using a network quality assessment algorithm and method. Finally, it combines video similarity, audio similarity, text similarity, network quality assessment, and time series evaluation to form a comprehensive audio and video perception evaluation, resulting in a user perception evaluation. This method addresses both the poor repeatability of subjective evaluation methods and the inability of objective methods to replicate human thought patterns, making it closer to human thinking and more aligned with users' perception of network audio and video call quality. Moreover, by mapping time and location, combined with network parameter information and events, it can more accurately pinpoint network problems.
[0015] Preferably, the specific steps of step S1 are as follows:
[0016] S11: The sending end initiates video data, and at the same time as the video data is initiated, it records the time series data in the process, and uploads the recorded time series data to the transceiver time series storage module of the server through the communication network for storage;
[0017] S12: The sending end initiates audio data and processes the audio data, that is, converts the audio into sending end text information, and records the time series data and network data during the process; the converted receiving end text information is uploaded to the server's text information storage module for storage; the recorded time series data is uploaded to the server's time series storage module for storage; and the recorded network data is uploaded to the server's network data storage module for storage.
[0018] S13: Simultaneously, the video and audio data sent by the sending end are saved through the storage module.
[0019] Preferably, the specific steps of step S2 are as follows:
[0020] S21: The receiving end initiates video data, and while receiving the video data, it records the time series data during the process, and uploads the recorded time series data to the transceiver time series storage module of the server through the communication network for storage;
[0021] S22: The receiving end initiates audio data and processes the audio data by converting the audio into receiving end text information, while recording the time series data and network data during the process; the converted receiving end text information is uploaded to the server's text information storage module for storage; the recorded time series data is uploaded to the server's time series storage module for storage; and the recorded network data is uploaded to the server's network data storage module for storage.
[0022] S23: Simultaneously, the received video and audio data are saved through the storage module.
[0023] Preferably, the specific steps of step S3 are as follows:
[0024] S31: Use video similarity to evaluate the similarity between the video data from the sending end in step S1 and the video data from the receiving end in step S2, and display it in real time;
[0025] S32: Use text similarity to evaluate the similarity between the text information data converted from audio data processing at the sending end in step S1 and the text information data converted from audio data processing at the receiving end in step S2, and display the results in real time.
[0026] S33: Use voice information to establish a user perception evaluation model through telecommunications psychology algorithms to evaluate users' voice perception;
[0027] S34: Network quality is evaluated based on network parameters and event information from the sending and receiving ends using network quality evaluation algorithms and methods;
[0028] S35: Combining steps S31, S32, S33 and S34, a comprehensive evaluation of audio and video perception is conducted to ultimately form a user perception evaluation.
[0029] Preferably, the specific steps for evaluating video similarity using video similarity in step S31 are as follows:
[0030] S311: The sending end divides the original video data according to the time stored in the time series data, forming a correlation between each frame of the original image and a time point;
[0031] S312: The receiving end collects the comparison video of the original video data through the communication network, and divides the comparison video into frames according to the time of the video frames stored in the time series, forming a correspondence between each frame comparison image and the comparison time point;
[0032] S313: Calculate the image similarity between each original image at each time point and each frame of the comparison image at the comparison time point. Based on the image similarity calculation results, evaluate the video quality and finally output the results.
[0033] Preferably, the specific steps for evaluating file similarity using text similarity in step S32 are as follows:
[0034] S321: Generate a corresponding standard audio segment as a comparison audio by mechanically reading aloud the original audio data from the sending end, and then convert the comparison audio into the original text;
[0035] S322: The receiving end (another terminal or platform) collects the comparison audio through a communication network and then converts it into comparison text;
[0036] S323: Calculate the text similarity between the original text and the comparison text using a text similarity algorithm, then transform them through function mapping, and finally output the result.
[0037] Preferably, in step S33, the speech perception evaluation using telecommunications psychology algorithms involves various speech samples undergoing manual perception evaluation to establish a user speech perception evaluation model. The specific steps for evaluating speech perception are as follows:
[0038] S331 Data Acquisition: Collect voice audio files and corresponding VoLTE / VONR network indicators from the sender and receiver under different network quality conditions, including call setup delay, jitter, voice packet loss rate, IP packet delay, and handover interruption delay;
[0039] S332 Data Processing: Users listen to the audio files from both the voice initiator and the voice receiver, and vote on the quality of the audio based on their personal perception. A threshold is set based on the voting results. If a user gives a good score exceeding the threshold, the audio file is labeled with tag 1; tag 0 indicates a bad score given by a user exceeding the threshold. Thus, each VoLTE / VONR network indicator has its corresponding perception tag.
[0040] S333 Feature Selection: Before building the classification model, feature scores in XGBoost are used to screen the final variables; feature variables are screened to prevent some variables from being too highly correlated;
[0041] S334 Model Establishment: Based on the existing network metrics corresponding to good and bad audio, multiple classification algorithms are used to train the training set, and the test set is used for verification to obtain the optimal classification model and output the user perception model.
[0042] S335 Model Prediction: Perform user perception model prediction on the corresponding network indicators of the audio and map the perception probability to the user perception score; in step S334, the score table of each audio file can be output through the established classification model.
[0043] Preferably, the specific steps of the network quality evaluation algorithm and method in step S34 are as follows:
[0044] S341 Data Collection: Collects network data including network parameters and event information, namely user GPS information, MR data and VoLTE / VONR data;
[0045] S342 Data Processing: Integrate and associate the data sources in step S322 at the raster level;
[0046] S343 Data Calculation and Analysis: Before calculating the performance indicators of the grid network, the basic network performance score of each cell in the grid is calculated first. After obtaining the basic network performance scores of all cells in the grid, the basic network performance score of the grid is obtained by using an algorithm.
[0047] S344 data analysis results: The service type is VoLTE / VONR service. The network performance score for each grid is obtained by selecting the time period to be evaluated. The grid network performance score is divided into five ranges: Excellent, Good, Average, Poor, and Severe.
[0048] The threshold values of each indicator are adjusted to accurately reflect the current network quality. In particular, it enables VoLTE / VONR network performance evaluation for 50*50 grids, which is more in line with the needs of mobile network optimization. By utilizing the correlation and constraint relationships between indicator sets, it enables reasonable and objective evaluation of the network quality of micro-areas (50*50 grids, hereinafter referred to as grids), effectively guiding network optimization.
[0049] Preferably, the basic network performance score of the cell in step S343 The scores of all call statistics indicators (KPIs), i.e. The weighted sum is used to calculate the basic network performance score for each cell in the coverage grid. That is, the call statistics KPI score for each cell is calculated using different algorithms based on the indicator attributes, specifically:
[0050] If a smaller indicator is better, then When the time is right, the calculation formula is:
[0051] in, KPIs for all communities j The value of the indicator within the 2.5%-97.5% quantile range, KPIs for Community X j The range of intervals, where the numerator is the KPI in cell X. j The cumulative distribution function (AUC), with KPI as the denominator. j The value corresponding to the cell with the largest cumulative distribution function;
[0052] If the KPI of community X j Less than When the left endpoint is reached, the calculation formula is:
[0053] If the KPI of community X j Greater than When the right endpoint is reached, the calculation formula is:
[0054] If a higher indicator is always better, then... When the time is right, the calculation formula is:
[0055] If the KPI of community X j Greater than The right endpoint is then calculated using the following formula:
[0056] If the KPI of community X j Less than The formula for calculating the left endpoint is:
[0057] The final result is the basic network performance score covering all cells in the grid;
[0058] In step S343, after obtaining the basic network performance scores of all cells covering the grid in the data calculation and analysis, the basic network performance score of the grid is obtained by using an algorithm. The algorithm in the text is as follows:
[0059] Among them, Grid X Referring to a specific grid cell X, Refers to the set of all cells covering grid X;
[0060] After obtaining the basic performance score of the raster based on the above algorithm logic, user-based MR data within the raster is then added as an adjustment parameter. Thus, the final basic network performance score for each grid is obtained, using the following formula:
[0061] The range of this adjustment parameter is: in The value of raster X is the normalized value of the 14-day RSRP mean for all rasters. For each raster, the mean SINR over a 14-day period is used to normalize the mean SINR of the raster by performing a min-max normalization.
[0062] Normalization of the min-max data, also known as deviation standardization, is a linear transformation of the original data, mapping the result to the range of 0-1. The transformation function is:
[0063] Where max is the maximum value of the sample data and min is the minimum value of the sample data;
[0064] Finally, the basic network performance score of the raster is obtained based on the basic network performance score of the raster and the adjustment parameters. The formula is:
[0065] Finally, the basic network performance score is calculated. The score is mapped to the interval (0, 100).
[0066] Preferably, the method for comprehensive evaluation of audio and video perception in step S35 specifically includes the following steps:
[0067] After obtaining three user speech perception scores—video perception evaluation, speech perception evaluation, network quality evaluation, and text similarity—different weights were assigned to the results of the three methods based on experience, and a weighted average was used to obtain the final user speech perception score. The weights for the video perception evaluation method are SW0, speech perception evaluation method is SW1, network quality evaluation method is SW2, and text similarity method is SW3. The final comprehensive user speech perception evaluation is calculated using the following formula: S ensemble =SW0*S0+SW1*S1+SW2*S2+SW3*S3;
[0068] Wherein: S0 is the score result based on the video perception evaluation method, S1 is the score result based on the speech perception evaluation method, S2 is the score result based on the network quality evaluation method, and S3 is the score result based on the text similarity method.
[0069] Preferably, four text similarity algorithms are used in step S323. The text similarity algorithms in step S33 include four text similarity algorithms, namely: a statistical algorithm based on term frequency (TF), a Simhash text similarity algorithm, a text similarity algorithm based on vector space model (VSM), and a text similarity algorithm based on LDA topic model.
[0070] The specific steps of the statistical algorithm based on term frequency (TF) are as follows:
[0071] S332-1-1: List the individual words in the standard text;
[0072] S332-1-2: Calculate the frequency (f4, f5) of each character in the standard text in both the standard text and the comparison text. F , where f F ≤f4, extraneous text in the comparison text is not included in the statistics;
[0073] S332-1-3: Text similarity results, the formula is: Similarity tf The similarity of term frequencies (TFs);
[0074] The specific steps of the Simhash text similarity algorithm include:
[0075] S332-2-1: Segment the text into words and take the features and weights of the top n words with the highest TF-IDF weights; that is, a text becomes a set of (feature: weight) of length n.
[0076] S332-2-2: After performing a regular hash on the feature of the word, a 64-bit binary number is obtained, resulting in a set of length 20 (hash: weight);
[0077] S332-2-3: Based on the binary number hash obtained in step S332-2-2, the corresponding positions are 1 or 0. Take the positive value weight and negative value weight for the corresponding positions. For example, a word obtained after step S332-2-2 is (010111:5). After step S332-2-3, a list [-5,5,-5,5,5,5] can be obtained. That is, for a document, a list of 20 elements of length 64 is obtained [weight,-weight…weight].
[0078] S332-2-4: Add column vectors to the n lists in step S332-2-3 to obtain a single list; for example, [-5,5,-5,5,5,5], [-3,-3,-3,3,-3,3], and [1,-1,-1,1,1,1] are added together to obtain [-7,1,-9,9,3,9]. This results in a list of length 64 for a document.
[0079] S332-2-5: Judge each value in the list obtained in step S332-2-4. When it is negative, take 0, and when it is positive, take 1. For example, [-7,1,-9,9,3,9] results in 010111, thus obtaining a list of length 64 for a text.
[0080] S332-2-6: Calculate similarity by taking the XOR of the simhash values of the two texts. A result of 1 indicates they are different, and 0 indicates they are the same. The length of the text with a result of 1 divided by the overall length is the difference score. Subtracting the difference score from 1 gives the text similarity score. 4iST24T ;
[0081] The basic idea of VSM is to simplify text into an N-dimensional vector representation with the weights of feature terms (keywords). The model assumes that words are unrelated and uses vectors to represent text, thus simplifying the complex relationships between keywords. The text is represented by very simple vectors, making the model computationally achievable. Here, D stands for Document, and T stands for Term, representing a feature term. A feature term refers to the basic linguistic unit present in document D that can represent the content of that document; it is mainly composed of words or phrases. Text can be represented by a feature term set as D(T1, T2, ..., T). n ), where T k These are feature terms, requiring 1 <= k <= N; the specific steps of the text similarity algorithm based on the Vector Space Model (VSM) include:
[0082] S332-3-1: Suppose a speech text has four feature terms a, b, c, and d, then this speech text is represented as D(a, b, c, d);
[0083] S332-3-2: For other texts to be compared with the spoken text, the order of these feature terms is also followed; for a text containing n feature terms, each feature term is assigned a certain weight to represent its importance, i.e., D = D(T1, W1; T2, W2; ..., T...). n W n This can be abbreviated as D = D(W1, W2, ..., W...). n ), is called the weight vector representation of text D; where W k It is Tk The weights are 1 <= k <= N;
[0084] S332-3-3: In the vector space model, the content relevance Sim(D1,D2) between texts D1 and D2 is represented by the cosine of the angle between the vectors, and the formula is:
[0085] The specific steps of the text similarity algorithm based on the LDA topic model include: firstly, modeling the text set using the LDA model, that is, using the statistical characteristics of the text to map the text corpus to various topic spaces, mining the relationships between different topics and words hidden in the text, obtaining the topic distribution of the text, and then calculating the text similarity matrix through the topic distribution; the LDA model is a probabilistic topic model for modeling discrete datasets (such as document sets), and is a method for modeling the topic information of text data. By providing a brief description of the text, it retains the essential statistical information, which helps to efficiently process large-scale document sets;
[0086] The process of generating text using the LDA topic probability model is as follows:
[0087] S332-4-1: For topic z, a word multinomial distribution vector on that topic is obtained according to the Dirichlet distribution Dir(β).
[0088] S332-4-2: Obtain the number of words N in the text based on the Poisson distribution P;
[0089] S332-4-3: Obtain a topic distribution probability vector θ for this text based on the Dirichlet distribution Dir(α);
[0090] S332-4-4: For each word Wn in the N words of this text:
[0091] S332-4-5: Randomly select a topic z from the multinomial(θ) distribution of θ;
[0092] S332-4-6: Select a word as Wn from the multinomial(Φ) conditional probability distribution of topic z;
[0093] Since the topic distribution of text is a simple mapping of the text vector space, calculating the similarity between two texts in the context of topic representation involves calculating the corresponding topic probability distribution. Because topics are a mixed distribution of word vectors, the Kullback-Leibler relative entropy distance, denoted as KL distance, is used as the similarity metric. The formula for calculating KL distance is: Where DKV (p,q) represents the information loss that occurs when the probability distribution Q is used to fit the true distribution P, where P represents the true distribution and q represents the fitted distribution of P.
[0094] Preferably, the ViNR-based user audio and video perception evaluation system includes a transmitting module, a receiving module, and a perception evaluation module. The transmitting module and the receiving module are connected via a communication network, and both the transmitting module and the receiving module are communicatively connected to the perception evaluation module. The transmitting module includes a transmitting audio and video data unit and a communication unit one, which forms data connections with the transmitting audio and video data unit and the transmitting audio and video data processing unit, respectively. The receiving module includes a receiving audio and video data unit and a communication unit two, which forms data connections with the receiving audio and video data unit and the receiving audio and video data processing unit, respectively. The perception evaluation module includes a transmitting audio and video data processing unit, a data time series processing unit, a receiving audio and video data processing unit, a video storage unit, a video similarity evaluation unit, an audio storage unit, an audio similarity evaluation unit, and a time series storage unit. The system comprises a time-series evaluation unit, a text storage unit, a text similarity evaluation unit, a signal storage unit, a network quality evaluation unit, and a user perception evaluation unit. Both the transmitting and receiving audio-visual data processing units are communicatively connected to these units. The video storage unit is electrically connected to the video similarity evaluation unit, the audio storage unit is electrically connected to the audio similarity evaluation unit, the time-series processing unit is electrically connected to the time-series storage unit, the time-series storage unit is electrically connected to the time-series evaluation unit, the text storage unit is electrically connected to the text similarity evaluation unit, and the signal storage unit is electrically connected to the network quality evaluation unit. The video similarity evaluation unit, audio similarity evaluation unit, time-series evaluation unit, text similarity evaluation unit, and network quality evaluation unit are all electrically connected to the user perception evaluation unit. The transmitting module communicates with the transmitting audio-visual data processing unit of the perception evaluation module via communication unit one, and the receiving module communicates with the receiving audio-visual data processing unit of the perception evaluation module via communication unit two.
[0095] The above technical solution includes a video storage unit for storing video information processed by the transmitting and receiving audio-visual data processing units, an audio storage unit for storing audio information processed by the transmitting and receiving audio-visual data processing units, a time-series storage unit for storing time-series information processed by the transmitting and receiving audio-visual data processing units, a text storage unit for storing text information processed by the transmitting and receiving audio-visual data processing units, and a signal storage module for storing network parameters and event information of the transmitting and receiving modules. The audio-visual transmitting module, audio-visual receiving module, and perception evaluation module are combined to form a network user audio-visual perception evaluation system, thereby achieving audio-visual perception evaluation of network users.
[0096] Compared with existing technologies, the beneficial effects of this invention are as follows: by using video similarity, audio similarity, and text similarity algorithms to judge the user's perceived audio and video quality, it not only solves the problem of poor repeatability of the MOS subjective evaluation method, but also solves the problem that the objective problem of MOS-LQO cannot reproduce the human brain's thinking paradigm. It is closer to the human brain's thinking mode and closer to the user's perception of the audio and video quality of network calls. At the same time, by mapping time and location, combined with network parameters and events, network problems can be located more accurately. Attached Figure Description
[0097] Figure 1 is a schematic diagram of the ViNR user audio and video perception evaluation method and process of the present invention.
[0098] Figure 2 is a schematic diagram of the video similarity evaluation process of the ViNR user audio and video perception evaluation method of the present invention;
[0099] Figure 3 is a flowchart illustrating the text similarity evaluation process of the ViNR user audio and video perception evaluation method of the present invention.
[0100] Figure 4 is a system framework diagram of the ViNR user audio and video perception evaluation method of the present invention. Detailed Implementation
[0101] The technical solutions in the embodiments of the present invention will now be clearly and completely described with reference to the accompanying drawings.
[0102] Example: As shown in Figure 1, the ViNR user audio and video perception evaluation method specifically includes the following steps:
[0103] S1 Audio and Video Data Initiation: The sending end initiates audio and video data, records and saves the time-series data corresponding to the audio and video data at the time of initiation; the audio data is processed and saved, and the network data of the sending end is also saved.
[0104] The specific steps of step S1 are as follows:
[0105] S11: The sending end initiates video data, and at the same time as the video data is initiated, it records the time series data in the process, and uploads the recorded time series data to the transceiver time series storage module of the server through the communication network for storage;
[0106] S12: The sending end initiates audio data and processes the audio data, that is, converts the audio into sending end text information, and records the time series data and network data during the process; the converted receiving end text information is uploaded to the server's text information storage module for storage through the communication network; the recorded time series data is uploaded to the server's time series storage module for storage through the communication network; and the recorded network data is uploaded to the server's network data storage module for storage through the communication network.
[0107] S13: Simultaneously, the video and audio data sent by the sending end are saved through the storage module;
[0108] S2 Audio and Video Data Reception: The receiving end receives audio and video data, processes and saves the time sequence corresponding to the received audio and video data; in particular, the audio data is processed and saved, and the network data of the receiving end is also saved.
[0109] The specific steps of step S2 are as follows:
[0110] S21: The receiving end initiates video data, and while receiving the video data, it records the time series data during the process, and uploads the recorded time series data to the transceiver time series storage module of the server through the communication network for storage;
[0111] S22: The receiving end initiates audio data and processes the audio data, that is, converts the audio into receiving end text information, and records the time series data and network data during the process; the converted receiving end text information is uploaded to the server's text information storage module for storage through the communication network; the recorded time series data is uploaded to the server's time series storage module for storage through the communication network; and the recorded network data is uploaded to the server's network data storage module for storage through the communication network.
[0112] S23: Simultaneously, the received video and audio data from the receiving end are saved through the storage module;
[0113] S3 Evaluation and Analysis: Similarity evaluation methods are used to evaluate various data from both the sending and receiving ends. Time series data is combined for time series evaluation, and network data is combined for network quality evaluation. Finally, a comprehensive audio and video perception evaluation is obtained to form a user perception evaluation.
[0114] The specific steps of step S3 are as follows:
[0115] S31: Use video similarity to evaluate the similarity between the video data from the sending end in step S1 and the video data from the receiving end in step S2, and display it in real time;
[0116] As shown in Figure 2, the specific steps for evaluating video similarity using video similarity in step S31 are as follows:
[0117] S311: The sending end divides the original video data according to the time stored in the time series data, forming a correlation between each frame of the original image and a time point;
[0118] S312: The receiving end (i.e., using a different terminal or platform than the sending end) collects the comparison video of the original video data through the communication network, and divides the comparison video into frames according to the time of the video frames stored in the time series, forming a correspondence between each frame comparison image and the comparison time point.
[0119] S313: Calculate the image similarity between each frame of the original image at each time point and each frame of the comparison image at the comparison time point. Based on the image similarity calculation results, evaluate the video quality and finally output the results. To more accurately evaluate video quality and better reflect actual user perception, we selected the image similarity method for video quality evaluation. The invention relates to an algorithm framework for calculating image similarity, and the specific steps of step S313 are as follows:
[0120] S3131: Preprocess the image, scale the image, and adjust the image to a uniform size to facilitate subsequent feature extraction and comparison;
[0121] S3132: Perform grayscale processing to convert the color image into a grayscale image, reducing the impact of color on similarity calculation and lowering computational complexity;
[0122] S3133: Normalize the image to eliminate the influence of factors such as lighting on image similarity;
[0123] S3134: SIFT extracts keypoints or feature vectors from an image. These features should reflect the main information of the image, such as shape, texture, and color. Depending on the specific algorithm, a descriptor or feature vector is calculated for each keypoint.
[0124] In S3134, the SIFT (Scale-Invariant Feature Transform) algorithm for extracting key points from images is mentioned. This algorithm is a feature extraction algorithm used in image processing and computer vision. The specific steps for using the SIFT algorithm are as follows:
[0125] S31341 Scale Space Pyramid: By using Gaussian blur functions of different scales to filter the original image, a series of image pyramids are generated, also known as scale space;
[0126] S31342 Keypoint Localization: Detects potential keypoint locations using the Difference of Gaussian (DoG) scale space; keypoints are stable in both space and scale, maintaining certain invariance under different scales and rotations;
[0127] S31343 Keypoint Orientation Determination: After determining the potential keypoint locations, the SIFT algorithm calculates the gradient magnitude and direction in the neighborhood around each keypoint and uses the gradient histogram to determine the main orientation of the keypoint.
[0128] S31344 Keypoint Descriptor Generation: Based on the location, scale, and orientation information of keypoints, a descriptor vector is generated for each keypoint, containing detailed information about the image region surrounding the keypoint for subsequent feature matching.
[0129] S31345 Feature Matching: Using the generated descriptors, search for feature points in the target image that are similar to key points in the reference image;
[0130] S3135 Identify key points: Compare the features of the image and the original image to find the key points that are similar to each other;
[0131] S3136 Calculate similarity: Calculate a similarity score based on the matched feature points, such as based on a distance metric (Euclidean distance) or a similarity metric (cosine similarity, etc.). Step S336 mentions using a similarity metric (cosine similarity). Cosine similarity is a commonly used method to calculate the similarity between two vectors, based on the cosine of the angle between them. The smaller the angle between two vectors, the larger their cosine value, indicating that the two vectors are more similar. Specifically:
[0132] S31361: First, obtain the descriptor vectors of the original image and the comparison image, and set them as vector A and vector B respectively;
[0133] S31362: Calculate cosine similarity. The formula for calculating cosine similarity is:
[0134] S31363: To calculate the dot product, multiply the corresponding elements of vectors A and B, and then add all the products together.
[0135] S31364: Sum the squares of each element in the vector, then take the square root. Use the above formula to calculate the cosine similarity;
[0136] S3137: Calculate the similarity between the two images based on the feature matching results. For missing comparison images, set the similarity to 0;
[0137] S3138: Accumulate the similarity of each frame and calculate the average as the final video similarity result;
[0138] S32: Use text similarity to evaluate the similarity between the text information data converted from audio data processing at the sending end in step S1 and the text information data converted from audio data processing at the receiving end in step S2, and display the results in real time.
[0139] As shown in Figure 3, the specific steps for evaluating file similarity using text similarity in step S32 are as follows:
[0140] S321: Generate a corresponding standard audio segment as a comparison audio by mechanically reading aloud the original audio data from the sending end, and then convert the comparison audio into the original text;
[0141] S322: The receiving end (another terminal or platform) collects the comparison audio through a communication network and then converts it into comparison text;
[0142] S323: Calculate the text similarity between the original text and the comparison text using a text similarity algorithm, then transform them through function mapping, and finally output the result;
[0143] In step S323, four text similarity algorithms are used. The text similarity algorithms in step S33 include four text similarity algorithms, namely: a statistical algorithm based on term frequency (TF), a Simhash text similarity algorithm, a text similarity algorithm based on vector space model VSM, and a text similarity algorithm based on LDA topic model.
[0144] The specific steps of the statistical algorithm based on term frequency (TF) are as follows:
[0145] S332-1-1: List the individual words in the standard text;
[0146] S332-1-2: Calculate the frequency (f4, f5) of each character in the standard text in both the standard text and the comparison text. F , where f F ≤f4, extraneous text in the comparison text is not included in the statistics;
[0147] S332-1-3: Text similarity results, the formula is: Similarity tf The similarity of term frequencies (TFs);
[0148] The specific steps of the Simhash text similarity algorithm include:
[0149] S332-2-1: Segment the text into words and take the features and weights of the top n words with the highest TF-IDF weights; that is, a text becomes a set of (feature: weight) of length n.
[0150] S332-2-2: After performing a regular hash on the feature of the word, a 64-bit binary number is obtained, resulting in a set of length 20 (hash: weight);
[0151] S332-2-3: Based on the binary number hash obtained in step S332-2-2, the corresponding positions are 1 or 0. Take the positive value weight and negative value weight for the corresponding positions. For example, a word obtained after step S332-2-2 is (010111:5). After step S332-2-3, a list [-5,5,-5,5,5,5] can be obtained. That is, for a document, a list of 20 elements of length 64 is obtained [weight,-weight…weight].
[0152] S332-2-4: Add column vectors to the n lists in step S332-2-3 to obtain a single list; for example, [-5,5,-5,5,5,5], [-3,-3,-3,3,-3,3], and [1,-1,-1,1,1,1] are added together to obtain [-7,1,-9,9,3,9]. This results in a list of length 64 for a document.
[0153] S332-2-5: Judge each value in the list obtained in step S332-2-4. When it is negative, take 0, and when it is positive, take 1. For example, [-7,1,-9,9,3,9] results in 010111, thus obtaining a list of length 64 for a text.
[0154] S332-2-6: Calculate similarity by taking the XOR of the simhash values of the two texts. A result of 1 indicates they are different, and 0 indicates they are the same. The length of the text with a result of 1 divided by the overall length is the difference score. Subtracting the difference score from 1 gives the text similarity score. 4iST24T ;
[0155] The basic idea of VSM is to simplify text into an N-dimensional vector representation with the weights of feature terms (keywords). The model assumes that words are unrelated and uses vectors to represent text, thus simplifying the complex relationships between keywords. The text is represented by very simple vectors, making the model computationally achievable. Here, D stands for Document, and T stands for Term, representing feature terms. A feature term refers to a basic linguistic unit appearing in document D and representing its content; it is mainly composed of words or phrases. Text can be represented by a feature term set as D(T1, T2, ..., T). n ), where T k These are feature terms, requiring 1 <= k <= N; the specific steps of the text similarity algorithm based on the Vector Space Model (VSM) include:
[0156] S332-3-1: Suppose a speech text has four feature terms a, b, c, and d, then this speech text is represented as D(a, b, c, d);
[0157] S332-3-2: For other texts to be compared with the spoken text, the order of these feature terms is also followed; for a text containing n feature terms, each feature term is assigned a certain weight to represent its importance, i.e., D = D(T1, W1; T2, W2; ..., T...). n W n This can be abbreviated as D = D(W1, W2, ..., W...). n ), is called the weight vector representation of text D; where W k It is T k The weights are 1 <= k <= N;
[0158] S332-3-3: In the vector space model, the content relevance Sim(D1,D2) between texts D1 and D2 is represented by the cosine of the angle between the vectors, and the formula is:
[0159] The specific steps of the text similarity algorithm based on the LDA topic model include: firstly, modeling the text set using the LDA model, that is, using the statistical characteristics of the text to map the text corpus to various topic spaces, mining the relationships between different topics and words hidden in the text, obtaining the topic distribution of the text, and then calculating the text similarity matrix through the topic distribution; the LDA model is a probabilistic topic model for modeling discrete datasets (such as document sets), and is a method for modeling the topic information of text data. By providing a brief description of the text, it retains the essential statistical information, which helps to efficiently process large-scale document sets;
[0160] The process of generating text using the LDA topic probability model is as follows:
[0161] S332-4-1: For topic z, a word multinomial distribution vector on that topic is obtained according to the Dirichlet distribution Dir(β).
[0162] S332-4-2: Obtain the number of words N in the text based on the Poisson distribution P;
[0163] S332-4-3: Obtain a topic distribution probability vector θ for this text based on the Dirichlet distribution Dir(α);
[0164] S332-4-4: For each word Wn in the N words of this text:
[0165] S332-4-5: Randomly select a topic z from the multinomial(θ) distribution of θ;
[0166] S332-4-6: Select a word as Wn from the multinomial(Φ) conditional probability distribution of topic z;
[0167] Since the topic distribution of text is a simple mapping of the text vector space, calculating the similarity between two texts in the context of topic representation involves calculating the corresponding topic probability distribution. Because topics are a mixed distribution of word vectors, the Kullback-Leibler relative entropy distance, denoted as KL distance, is used as the similarity metric. The formula for calculating KL distance is: Where D KV (p,q) represents the information loss that occurs when the probability distribution Q is used to fit the true distribution P, where P represents the true distribution and q represents the fitted distribution of P.
[0168] S33: Use voice information to establish a user perception evaluation model through telecommunications psychology algorithms to evaluate users' voice perception;
[0169] In step S33, the speech perception evaluation is performed using telecommunications psychology algorithms. This involves evaluating various speech samples through manual perception to establish a user speech perception evaluation model. The specific steps for speech perception evaluation are as follows:
[0170] S331 Data Acquisition: Collect voice audio files and corresponding VoLTE / VONR network indicators from the sender and receiver under different network quality conditions, including call setup delay, jitter, voice packet loss rate, IP packet delay, and handover interruption delay;
[0171] S332 Data Processing: Users listen to the audio files from both the voice initiator and the voice receiver, and vote on the quality of the audio based on their personal perception. A threshold is set based on the voting results. If a user gives a good score exceeding the threshold, the audio file is labeled with tag 1; tag 0 indicates a bad score given by a user exceeding the threshold. Thus, each VoLTE / VONR network indicator has its corresponding perception tag.
[0172] S333 Feature Selection: Before building the classification model, feature scores in XGBoost are used to screen the final variables; feature variables are screened to prevent some variables from being too highly correlated;
[0173] S334 Model Establishment: Based on the existing network metrics corresponding to good and bad audio, multiple classification algorithms are used to train the training set, and the test set is used for verification to obtain the optimal classification model and output the user perception model.
[0174] The multiple classification algorithms mentioned in step S334 include four classification algorithms: 1) Decision Tree; 2) Random Forest; 3) Logistic Regression; 4) XGBoost algorithm; among which,
[0175] 1) The specific steps of the decision tree algorithm are as follows:
[0176] S334-1-1: Select an optimal predictor variable to divide all sample units into two classes, maximizing the purity of both classes; if the predictor variable is continuous, select a split point for classification to maximize the purity of both classes; if the predictor variable is a categorical variable, merge and reclassify the categories.
[0177] S334-1-2: Continue performing the steps in step S334-1-1 for each subcategory;
[0178] S334-1-3: Repeat steps S334-1-1 to S334-1-2 until the number of sample units in the subclass is too small, or no classification method can reduce the impurity to below a given threshold; the final set of subclasses is the terminal node; determine the category of the terminal node based on the mode of the number of categories of sample units in each terminal node.
[0179] S334-1-4: Execute the decision tree for any sample unit to obtain its terminal node, that is, the category predicted by the model can be obtained according to step S334-1-3; however, usually a very large tree will be obtained by this algorithm, resulting in the phenomenon of overfitting and poor classification performance for units outside the training set; to solve the above problems, the 10-fold cross-validation method can be used to select the tree with the minimum prediction error;
[0180] 2) Random forest: A random forest is an ensemble classifier composed of a set of decision tree classifiers {h(X,θ k ), k = 1, 2, …, K}, where {θ k} is a random vector that follows an independent and identical distribution, and K represents the number of decision trees in the random forest. Given the independent variable X, each decision tree classifier determines the optimal classification result by voting; the random forest involves sampling the sample units and variables to generate a large number of decision trees; for each sample unit, all decision trees classify it in turn; the specific steps of the random forest algorithm are as follows:
[0181] S334-2-1: Apply the bootstrap method to randomly and with replacement draw K new bootstrap sample sets from the training set, and construct K classification trees from them. Each time, the samples not drawn form K out-of-bag data;
[0182] S334-2-2: Randomly draw m < M variables at each node of each tree. By calculating the information content contained in each variable, then select the variable with the most classification ability among the m variables for node splitting;
[0183] S334-2-3: Completely generate all decision trees without pruning;
[0184] S334-2-4: The category of the terminal node is determined by the mode category corresponding to the node;
[0185] S334-2-5: For a new observation point, classify it with all the trees, and its category is generated by the majority decision principle;
[0186] 3) The specific steps of the logistic regression algorithm are as follows:
[0187] S334-3-1 Establish a prediction function: First, construct a suitable prediction function, denoted as the h function. This function h is the classification function to be found. The output of this function h must be two values to predict the judgment result of the input data. Therefore, use the Logistic function, and the function form is:
[0188] Next, we need to determine the boundary type for data partitioning. Here, we only discuss the case of linear boundaries. For linear boundaries, the form is as follows:
[0189] Where θ represents the regression parameter and x represents the independent variable;
[0190] The prediction function is constructed as follows:
[0191] Where θ represents the regression parameter and x represents the independent variable;
[0192] h f The value of the function (x) represents the probability that the result is 1. Therefore, the probability that the classification result of input x is category 1 or category 0 is calculated according to the following formula: p(y|x;θ)=(h θ (x)) y (1-h θ (x)) 1-y y = 1, 0;
[0193] S33432 Establish the Cost function: any value h that can measure the model's prediction. f The difference function between (x) and the true value y is called the cost function. For each algorithm, the cost function is not unique. The following are common cross-entropy functions. After determining the function, the smaller cost function value J(θ) can be obtained by continuously changing the parameter θ.
[0194] Where m is the number of training samples, h f (x) represents the predicted value, and y represents the actual value;
[0195] 4) The specific steps of the XGBoost algorithm are as follows:
[0196] S334-4-1 defines the complexity of a tree: First, the tree is split into a structural part q and a leaf node weight part w, where w is a vector representing the output value of each leaf node, and T represents the number of leaf nodes in a decision tree; f t (x)=w q(x) ,w∈R T ,q:R d →{1,2,...,T};
[0197] Introducing the regularization term Ω(f) J This can be used to control the complexity of the tree, thereby effectively controlling the overfitting of the model;
[0198] Where T represents the number of leaf nodes in a decision tree, γ represents the coefficient that controls the complexity of the tree, which is equivalent to pre-pruning the tree of the XGBoost algorithm model, and λ represents the proportion by which the regularization term is changed, which is equivalent to penalizing the complex model to prevent the model from overfitting.
[0199] S334-4-2 Boosting Tree Model in XGBoost: Similar to the GBDT method, XGBoost's boosting model also uses residuals. However, the selection of split nodes is not necessarily based on least-squares loss. Its loss function is as follows, and compared to GBDT, it adds a regularization term based on the complexity of the tree model:
[0200] in, This represents the estimated value, y i Represents the actual value. Ω(f) represents the model residuals. _ This refers to the regularization term mentioned earlier;
[0201] S334-4-3 Rewrites the objective function: In XGBoost, the loss function is directly expanded into a binomial function using Taylor expansion, provided that the loss function is first-order, second-order, and continuously differentiable. Assume our leaf node region is: I j ={i|q(x i )=j};
[0202] Among them, I j ={i|q(x i )=j} represents the set of labels of the samples assigned to the j-th leaf node in the training samples. For example, if the 1st, 3rd, and 5th samples in the training samples are assigned to the 2nd leaf node, then j={1,3,5}.
[0203] For g i and h i They are defined as follows:
[0204] Among them, y i Represents the actual value. This represents the predicted value from iteration t-1;
[0205] The objective function for t trees can be transformed into the following by second-order Taylor expansion:
[0206] definition
[0207] At this point, w j Taking the derivative and setting it to 0, we get:
[0208] S334-4-4 Tree Structure Scoring Function: The Obj value above represents the maximum reduction in the target value when a given tree structure is specified; this can be called the structure score. It can be considered a more general scoring function for tree structures, similar to the Gini index. For finding the tree structure with the minimum Obj score, a greedy algorithm is used. Each time, it attempts to split the existing leaf nodes (the root node being the first leaf node), and then obtains the gain after the split:
[0209] The formula can be decomposed into the fraction on the left leaf, the fraction on the right leaf, the fraction on the original leaf, and the regularization on the additional leaf; here, Gain is used as the condition for determining whether to split.
[0210] If Gain < 0, this leaf node is not split; however, this still requires listing all splitting schemes for each split. In practice, all samples g are first... i Sort the data in ascending order, then iterate through the data to check if each node needs to be split. This splitting method only requires scanning the samples once to split the data into GL and GR, and then split the data according to the Gain score.
[0211] S335 Model Prediction: Perform user perception model prediction on the corresponding network indicators of the audio and map the perception probability to the user perception score; in step S334, the score table of each audio file can be output through the established classification model; that is, the score table of each audio file can be output through the established classification model in step S334.
[0212] S34: Network quality is evaluated based on network parameters and event information from the sending and receiving ends using a network quality evaluation algorithm and method. To evaluate user network quality by storing network parameters and event information from the voice initiator and receiver, the specific steps of the network quality evaluation algorithm and method in step S34 are as follows:
[0213] S341 Data Collection: Collects network data including network parameters and event information, namely user GPS information, MR data and VoLTE / VONR data;
[0214] S342 Data Processing: Integrate and associate the data sources in step S322 at the raster level;
[0215] S343 Data Calculation and Analysis: Before calculating the performance indicators of the grid network, the basic network performance score of each cell in the grid is calculated first. After obtaining the basic network performance scores of all cells in the grid, the basic network performance score of the grid is obtained by using an algorithm.
[0216] The basic network performance score of the cell in step S343 The scores of all call statistics indicators (KPIs), i.e. The weighted sum is used to calculate the basic network performance score for each cell in the coverage grid. That is, the call statistics KPI score for each cell is calculated using different algorithms based on the indicator attributes, specifically:
[0217] If the smaller the indicator, the better, when When the time is right, the calculation formula is:
[0218] in, KPIs for all communities j The value of the indicator within the 2.5%-97.5% quantile range, KPIs for Community X j The range of intervals, where the numerator is the KPI in cell X. j The cumulative distribution function (AUC), with KPI as the denominator. j The value corresponding to the cell with the largest cumulative distribution function;
[0219] If the KPI of community X j Less than When the left endpoint is reached, the calculation formula is:
[0220] If the KPI of community X j Greater than When the right endpoint is reached, the calculation formula is:
[0221] If a higher indicator is always better, then... When the time is right, the calculation formula is:
[0222] If the KPI of community X j Greater than The right endpoint is then calculated using the following formula:
[0223] If the KPI of community X j Less than The formula for calculating the left endpoint is:
[0224] The final result is the basic network performance score covering all cells in the grid;
[0225] In step S343, after obtaining the basic network performance scores of all cells covering the grid in the data calculation and analysis, the basic network performance score of the grid is obtained by using an algorithm. The algorithm in the text is as follows:
[0226] Among them, Grid X Referring to a specific grid cell X, Refers to the set of all cells covering grid X;
[0227] After obtaining the basic performance score of the raster based on the above algorithm logic, user-based MR data within the raster is then added as an adjustment parameter. Thus, the final basic network performance score for each grid is obtained, using the following formula:
[0228] The range of this adjustment parameter is: in The value of raster X is the normalized value of the 14-day RSRP mean for all rasters. For each raster, the mean SINR over a 14-day period is used to normalize the mean SINR of the raster by performing a min-max normalization.
[0229] Normalization of the min-max data, also known as deviation standardization, is a linear transformation of the original data, mapping the result to the range of 0-1. The transformation function is:
[0230] Where max is the maximum value of the sample data and min is the minimum value of the sample data;
[0231] Finally, the basic network performance score of the raster is obtained based on the basic network performance score of the raster and the adjustment parameters. The formula is:
[0232] Finally, the basic network performance score is calculated. The score is mapped to the interval (0, 100);
[0233] S344 data analysis results: The service type is VoLTE / VONR service. Then, select the time to be evaluated to obtain the network performance score of the grid. The network performance score of the grid is divided into 5 ranges: excellent, good, average, poor, and severe.
[0234] S35: Combining steps S31, S32, S33 and S34, a comprehensive evaluation of audio and video perception is conducted to ultimately form a user perception evaluation;
[0235] The method for comprehensive audio-visual perception evaluation in step S35 specifically includes the following steps:
[0236] After obtaining three user speech perception scores—video perception evaluation, speech perception evaluation, network quality evaluation, and text similarity—different weights were assigned to the results of the three methods based on experience, and a weighted average was used to obtain the final user speech perception score. The weights for the video perception evaluation method are SW0, speech perception evaluation method is SW1, network quality evaluation method is SW2, and text similarity method is SW3. The final comprehensive user speech perception evaluation is calculated using the following formula: S ensemble =SW0*S0+SW1*S1+SW2*S2+SW3*S3;
[0237] In this embodiment, the weights for the video perception evaluation method are 0.2, the speech perception evaluation method is 0.3, the network quality evaluation method is 0.2, and the text similarity method is 0.3; the final comprehensive user speech perception evaluation is calculated using the following formula: S ensemble =0.2*S0+0.3*S1+0.2*S2+0.3*S3;
[0238] Wherein: S0 is the score result based on the video perception evaluation method, S1 is the score result based on the speech perception evaluation method, S2 is the score result based on the network quality evaluation method, and S3 is the score result based on the text similarity method.
[0239] As shown in Figure 4, the ViNR-based user audio and video perception evaluation system includes a transmitting module, a receiving module, and a perception evaluation module. The transmitting module and the receiving module are connected via a communication network, and both the transmitting module and the receiving module are communicatively connected to the perception evaluation module. The transmitting module includes a transmitting audio and video data unit and a communication unit one, which forms data connections with the transmitting audio and video data unit and the transmitting audio and video data processing unit, respectively. The receiving module includes a receiving audio and video data unit and a communication unit two, which also forms data connections with the receiving audio and video data unit and the receiving audio and video data processing unit, respectively. The perception evaluation module includes a transmitting audio and video data processing unit, a data time series processing unit, a receiving audio and video data processing unit, a video storage unit, a video similarity evaluation unit, an audio storage unit, an audio similarity evaluation unit, and a time series storage unit. The system comprises a data time-series processing unit, a time-series evaluation unit, a text storage unit, a text similarity evaluation unit, a signal storage unit, a network quality evaluation unit, and a user perception evaluation unit. Both the transmitting and receiving audio-visual data processing units are communicatively connected to these units. The video storage unit is electrically connected to the video similarity evaluation unit, the audio storage unit is electrically connected to the audio similarity evaluation unit, the data time-series processing unit is electrically connected to the time-series storage unit, the time-series storage unit is electrically connected to the time-series evaluation unit, the text storage unit is electrically connected to the text similarity evaluation unit, the signal storage unit is electrically connected to the network quality evaluation unit, and the video similarity evaluation unit, audio similarity evaluation unit, time-series evaluation unit, text similarity evaluation unit, and network quality evaluation unit are all electrically connected to the user perception evaluation unit. The transmitting module communicates with the transmitting audio-visual data processing unit of the perception evaluation module via communication unit one, and the receiving module communicates with the receiving audio-visual data processing unit of the perception evaluation module via communication unit two.
[0240] Using the above technical solution, the overall audio and video data of the sending end is processed. The sending end video data is stored in the transceiver video storage unit, and the sending end video time-series data is also stored in the transceiver time-series storage unit. The sending end audio data is stored in the transceiver audio storage unit, and the sending end audio is converted into sending end text information, which is then stored in the transceiver text storage unit. The sending end text information time-series data is also stored in the transceiver time-series storage unit. Simultaneously, the sending end's network parameters and event information are saved. The sending end sends the audio and video data to the receiving end through the communication network. After receiving the audio and video data, the receiving end processes the overall audio and video data, storing the receiving end video data in the transceiver video storage unit, and the receiving end video time-series data in the transceiver time-series storage unit. The receiving end audio data is stored in the transceiver audio storage unit, and the receiving end audio is converted into receiving end text information, which is then stored in the transceiver text storage unit. The system is designed to perform the following steps: First, it stores the time-series text information from the receiving end into a transceiver time-series storage unit. Second, it saves the network parameters and event information from the receiving end. Third, it evaluates the similarity between the sending and receiving video data using video similarity methods and displays the results in real time. Fourth, it evaluates the similarity between the sending and receiving audio data using audio similarity methods and displays the results in real time. Fifth, it evaluates the similarity between the sending and receiving audio and text data using text similarity methods and displays the results in real time. Sixth, it evaluates the network quality based on the network parameters and event information from both the sending and receiving ends using network quality evaluation algorithms and methods. Finally, it evaluates the time-series data of the sending, receiving, and receiving ends using time-series alignment algorithms and methods. Finally, it combines video similarity, audio similarity, text similarity, network quality evaluation, and time-series evaluation to perform a comprehensive audio-visual perception evaluation, ultimately forming a user perception evaluation.
[0241] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A ViNR user audio and video perception evaluation method, characterized in that, Specifically, the following steps are included: S1 Audio and Video Data Initiation: The sending end initiates audio and video data, records and saves the time-series data corresponding to the audio and video data at the time of initiation; the audio data is processed and saved, and the network data of the sending end is also saved. S2 Audio and Video Data Reception: The receiving end receives audio and video data, processes and saves the time sequence corresponding to the received audio and video data; in particular, the audio data is processed and saved, and the network data of the receiving end is also saved. S3 Evaluation and Analysis: Similarity evaluation methods are used to evaluate various data from both the sending and receiving ends. Time series data is combined for time series evaluation, and network quality is combined for network data evaluation. Finally, a comprehensive audio-visual perception evaluation is obtained to form a user perception evaluation.
2. The ViNR user audio and video perception evaluation method according to claim 1, characterized in that, The specific steps of step S3 are as follows: The specific steps of step S3 are as follows: S31: Use video similarity to evaluate the similarity between the video data from the sending end in step S1 and the video data from the receiving end in step S2, and display it in real time; S32: Use text similarity to evaluate the similarity between the text information data converted from audio data processing at the sending end in step S1 and the text information data converted from audio data processing at the receiving end in step S2, and display the results in real time. S33: Use voice information to establish a user perception evaluation model through telecommunications psychology algorithms to evaluate users' voice perception; S34: Network quality is evaluated based on network parameters and event information from the sending and receiving ends using network quality evaluation algorithms and methods; S35: Combining steps S31, S32, S33 and S34, a comprehensive evaluation of audio and video perception is conducted to ultimately form a user perception evaluation.
3. The ViNR user audio and video perception evaluation method according to claim 2, characterized in that, The specific steps for evaluating video similarity using video similarity in step S31 are as follows: S311: The sending end divides the original video data according to the time stored in the time series data, forming a correlation between each frame of the original image and a time point; S312: The receiving end collects the comparison video of the original video data through the communication network, and divides the comparison video into frames according to the time of the video frames stored in the time series, forming a correspondence between each frame comparison image and the comparison time point; S313: Calculate the image similarity between each original image at each time point and each frame of the comparison image at the comparison time point. Based on the image similarity calculation results, evaluate the video quality and finally output the results.
4. The ViNR user audio and video perception evaluation method according to claim 3, characterized in that, The specific steps for evaluating file similarity using text similarity in step S32 are as follows: S321: Generate a corresponding standard audio segment as a comparison audio by mechanically reading aloud the original audio data from the sending end, and then convert the comparison audio into the original text; S322: The receiving end collects the comparison audio and then converts it into comparison text; S323: Calculate the text similarity between the original text and the comparison text using a text similarity algorithm, then transform them through function mapping, and finally output the result.
5. The ViNR user audio and video perception evaluation method according to claim 4, characterized in that, In step S33, the speech perception evaluation is performed using telecommunications psychology algorithms. This involves evaluating various speech samples through manual perception to establish a user speech perception evaluation model. The specific steps for speech perception evaluation are as follows: S331 Data Acquisition: Collect voice audio files from the sender and receiver under different network quality conditions and corresponding VoLTE / VONR network indicators; S332 Data Processing: Users listen to the audio files from both the voice initiator and the voice receiver, and vote on the quality of the audio based on their personal perception. A threshold is set based on the voting results. If a user gives a good score exceeding the threshold, the audio file is labeled with tag 1; tag 0 indicates a bad score given by a user exceeding the threshold. Thus, each VoLTE / VONR network indicator has its corresponding perception tag. S333 Feature Selection: Before building the classification model, feature scores in xgboost are used to screen the final variables; S334 Model Establishment: Based on the existing network metrics corresponding to good and bad audio, multiple classification algorithms are used to train the training set, and the test set is used for verification to obtain the optimal classification model and output the user perception model. S335 Model Prediction: Perform user perception model prediction on the corresponding network indicators of the audio and map the perception probability to the user perception score; in step S334, the score table of each audio file can be output through the established classification model.
6. The ViNR user audio and video perception evaluation method according to claim 5, characterized in that, The specific steps of the network quality evaluation algorithm and method in step S34 are as follows: S341 Data Collection: Collects network data including network parameters and event information, namely user GPS information, MR data and VoLTE / VONR data; S342 Data Processing: Integrate and associate the data sources in step S322 at the raster level; S343 Data Calculation and Analysis: Before calculating the performance indicators of the grid network, the basic network performance score of each cell in the grid is calculated first. After obtaining the basic network performance scores of all cells in the grid, the basic network performance score of the grid is obtained by using an algorithm. S344 data analysis results: The service type is VoLTE / VONR service. Then select the time to be evaluated to obtain the network performance score of the grid.
7. The ViNR user audio and video perception evaluation method according to claim 6, characterized in that, The basic network performance score of the cell in step S343 The scores of all call statistics indicators (KPIs), i.e. The weighted sum is used to calculate the basic network performance score for each cell in the coverage grid. That is, the call statistics KPI score for each cell is calculated using different algorithms based on the indicator attributes, specifically: If a smaller indicator is better: when When the time is right, the calculation formula is: in, KPIs for all communities j The value of the indicator within the 2.5%-97.5% quantile range, KPIs for Community X j The range of intervals, where the numerator is the KPI in cell X. j The cumulative distribution function, with KPI as the denominator. j The value corresponding to the cell with the largest cumulative distribution function; If the KPI of community X j Less than When the left endpoint is reached, the calculation formula is: If the KPI of community X j Greater than When the right endpoint is reached, the calculation formula is: If a larger indicator is always better: when When the time is right, the calculation formula is: If the KPI of community X j Greater than The right endpoint is then calculated using the following formula: If the KPI of community X j Less than The left endpoint is then calculated using the following formula: The final result is the basic network performance score covering all cells in the grid; In step S343, after obtaining the basic network performance scores of all cells covering the grid in the data calculation and analysis, the basic network performance score of the grid is obtained by using an algorithm. The algorithm in the text is as follows: Among them, Grid X Referring to a specific grid cell X, Refers to the set of all cells covering grid X; After obtaining the basic performance score of the raster based on the above algorithm logic, user-based MR data within the raster is then added as an adjustment parameter. Thus, the final basic network performance score for each grid is obtained, using the following formula: The range of this adjustment parameter is: in The value of raster X is the normalized value of the 14-day RSRP mean for all rasters. For each raster, the mean SINR over a 14-day period is used to normalize the mean SINR of the raster by performing a min-max normalization. Normalization of the min-max data, also known as deviation standardization, is a linear transformation of the original data, mapping the result to the range of 0-1. The transformation function is: Where max is the maximum value of the sample data and min is the minimum value of the sample data; Finally, the basic network performance score of the raster is obtained based on the basic network performance score of the raster and the adjustment parameters. The formula is: Finally, the basic network performance score is calculated. The score is mapped to the interval (0, 100).
8. The ViNR user audio and video perception evaluation method according to claim 7, characterized in that, The method for comprehensive audio-visual perception evaluation in step S35 specifically includes the following steps: After obtaining three user speech perception scores—video perception evaluation, speech perception evaluation, network quality evaluation, and text similarity—different weights were assigned to the results of the three methods, and a weighted average was used to obtain the final user speech perception score. The weights for the video perception evaluation method are SW0, speech perception evaluation method is SW1, network quality evaluation method is SW2, and text similarity method is SW3. The final comprehensive user speech perception evaluation is calculated using the following formula: S ensemble =SW0*S0+SW1*S1+SW2*S2+SW3*S3; Wherein: S0 is the score result based on the video perception evaluation method, S1 is the score result based on the speech perception evaluation method, S2 is the score result based on the network quality evaluation method, and S3 is the score result based on the text similarity method.
9. The ViNR user audio and video perception evaluation method according to claim 7, characterized in that, In step S323, four text similarity algorithms are used. The text similarity algorithms in step S33 include four text similarity algorithms, namely: a statistical algorithm based on term frequency (TF), a Simhash text similarity algorithm, a text similarity algorithm based on vector space model VSM, and a text similarity algorithm based on LDA topic model. The specific steps of the word frequency-based statistical algorithm are as follows: S332-1-1: List the individual words in the standard text; S332-1-2: Calculate the frequency (f4, f5) of each character in the standard text in both the standard text and the comparison text. F , where f F ≤f4, extraneous text in the comparison text is not included in the statistics; S332-1-3: Text similarity results, the formula is: Similarity tf The similarity of term frequencies (TFs); The specific steps of the Simhash text similarity algorithm include: S332-2-1: Segment the text into words and take the features and weights of the top n words with the highest TF-IDF weights; that is, a text becomes a set of (feature: weight) of length n. S332-2-2: After performing a regular hash on the feature of the word, a 64-bit binary number is obtained, resulting in a set of length 20 (hash: weight); S332-2-3: Based on the binary number hash obtained in step S332-2-2, the corresponding positions are 1 or 0. Take the positive value weight and negative value weight for the corresponding positions. S332-2-4: Add column vectors of the n lists in step S332-2-3 to obtain a single list; S332-2-5: Judge each value in the list obtained in step S332-2-4. When it is a negative value, take 0, and when it is a positive value, take 1. S332-2-6: Calculate similarity by taking the XOR of the simhash values of the two texts. A result of 1 indicates they are different, and 0 indicates they are the same. The length of the text with a value of 1 divided by the overall length is the difference score. Subtracting the difference score from 1 gives the text similarity score. 4iST24T ; The specific steps of the text similarity algorithm based on the Vector Space Model (VSM) include: S332-3-1: Suppose a speech text has four feature terms a, b, c, and d, then this speech text is represented as D(a, b, c, d); S332-3-2: For other texts to be compared with the spoken text, the order of these feature terms is also followed; for a text containing n feature terms, each feature term is assigned a certain weight to represent its importance, i.e., D = D(T1, W1; T2, W2; ..., T...). n W n This can be abbreviated as D = D(W1, W2, ..., W...). n ), is called the weight vector representation of text D; where Wk is the weight of Tk, 1<=k<=N; S332-3-3: In the vector space model, the content relevance Sim(D1,D2) between texts D1 and D2 is represented by the cosine of the angle between the vectors, and the formula is: The specific steps of the text similarity algorithm based on the LDA topic model include: firstly, modeling the text set using the LDA model, that is, using the statistical characteristics of the text to map the text corpus to various topic spaces, mining the relationship between different topics and words hidden in the text, obtaining the topic distribution of the text, and then calculating the text similarity matrix through the topic distribution. The process of generating text using the LDA topic probability model is as follows: S332-4-1: For topic z, a word multinomial distribution vector on that topic is obtained according to the Dirichlet distribution Dir(β). S332-4-2: Obtain the number of words N in the text based on the Poisson distribution P; S332-4-3: Obtain a topic distribution probability vector θ for this text based on the Dirichlet distribution Dir(α); S332-4-4: For each word Wn in the N words of this text: S332-4-5: Randomly select a topic z from the multinomial(θ) distribution of θ; S332-4-6: Select a word as Wn from the multinomial(Φ) conditional probability distribution of topic z; Since the topic distribution of text is a simple mapping of the text vector space, calculating the similarity between two texts in the context of topic representation involves calculating the corresponding topic probability distribution. Because topics are a mixed distribution of word vectors, the Kullback-Leibler relative entropy distance, denoted as KL distance, is used as the similarity metric. The formula for calculating KL distance is: Where D KV (p,q) represents the information loss that occurs when the probability distribution Q is used to fit the true distribution P, where P represents the true distribution and q represents the fitted distribution of P.
10. The ViNR user audio and video perception evaluation method according to any one of claims 1-9, characterized in that, This ViNR-based user audio and video perception evaluation system includes a transmitting module, a receiving module, and a perception evaluation module. The transmitting module and the receiving module are connected via a communication network, and both the transmitting module and the receiving module are communicatively connected to the perception evaluation module. The transmitting module includes a transmitting audio and video data unit and a first communication unit, which forms data connections with the transmitting audio and video data unit and the transmitting audio and video data processing unit, respectively. The receiving module includes a receiving audio and video data unit and a second communication unit, which forms data connections with the receiving audio and video data unit and the receiving audio and video data processing unit, respectively. The perception evaluation module includes a transmitting audio and video data processing unit, a data time series processing unit, a receiving audio and video data processing unit, a video storage unit, a video similarity evaluation unit, an audio storage unit, an audio similarity evaluation unit, and a time series storage unit. The system comprises a time-series evaluation unit, a text storage unit, a text similarity evaluation unit, a signal storage unit, a network quality evaluation unit, and a user perception evaluation unit. Both the transmitting and receiving audio-visual data processing units are communicatively connected to these units. The video storage unit is electrically connected to the video similarity evaluation unit, the audio storage unit is electrically connected to the audio similarity evaluation unit, the data time-series processing unit is electrically connected to the time-series storage unit, the time-series storage unit is electrically connected to the time-series evaluation unit, the text storage unit is electrically connected to the text similarity evaluation unit, and the signal storage unit is electrically connected to the network quality evaluation unit. The video similarity evaluation unit, audio similarity evaluation unit, time-series evaluation unit, text similarity evaluation unit, and network quality evaluation unit are all electrically connected to the user perception evaluation unit. The transmitting module communicates with the transmitting audio-visual data processing unit of the perception evaluation module via communication unit one, and the receiving module communicates with the receiving audio-visual data processing unit of the perception evaluation module via communication unit two.
Citation Information
Patent Citations
Method and device for completely evaluating 3G visual telephone quality
CN102158881A
Method and system for speech quality perception evaluation based on speech semantic recognition technology
CN108877839A
Video call quality evaluation method and device, electronic equipment and storage medium
CN109714557A
VINR user audio and video perception evaluation method
CN119172528A
SYSTEM AND METHOD FOR MACHINE LEARNING BASED QoE PREDICTION OF VOICE / VIDEO SERVICES IN WIRELESS NETWORKS
US20190073603A1