A method, device and equipment for ViNR user audio-video perception evaluation
By acquiring the time and area information of ViNR user audio and video perception requests, and utilizing background collection index data and target perception evaluation models, the accuracy and stability issues of ViNR user audio and video perception result evaluation in existing technologies have been resolved, thereby achieving rapid positioning and improved optimization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中国移动通信集团云南有限公司
- Filing Date
- 2026-02-25
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies rely on data alignment and synchronization between two ends when determining the audio and video perception results of ViNR users. This can easily lead to comparison errors due to clock drift, packet loss, or retransmission, affecting the accuracy and stability of the assessment.
By acquiring time information and area identifiers from audio and video perception requests, we determine background data collection metrics such as measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and uplink user plane rate. Using a pre-fitted target perception evaluation model, we determine the audio and video perception results.
It enables rapid location and retrospective analysis of perception problems in specific areas and time periods, improving troubleshooting and optimization efficiency, and achieving a comprehensive, objective, and accurate reflection of the true audio and video perception level of ViNR users.
Smart Images

Figure CN122160498A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of wireless technology, and in particular to a ViNR user audio and video perception evaluation method, apparatus, and device. Background Technology
[0002] ViNR, as the primary method for 5G network video calls, monitors audio and video quality in real time, and is increasingly valued by operators for providing users with high-quality audio and video.
[0003] Currently, the main method for determining audio and video perception results is to save the network data of the sending and receiving ends, and use similarity evaluation methods to evaluate various data from both ends to determine the audio and video perception results. This method relies on data alignment and synchronization between the two ends. Once clock drift, packet loss, or retransmission occurs, it is prone to comparison errors, resulting in issues with the accuracy and stability of the evaluation. Summary of the Invention
[0004] This disclosure provides a ViNR user audio and video perception assessment method, apparatus, and device to enable rapid localization and retrospective analysis of perception problems in specific areas and time periods, improve troubleshooting and optimization efficiency, and achieve the effect of comprehensively, objectively, and accurately reflecting the true audio and video perception level of ViNR users.
[0005] In a first aspect, embodiments of this disclosure provide a ViNR user audio and video perception assessment method, the method comprising: In response to receiving an audio / video perception request, the system obtains first time information and a first region identifier from the audio / video perception request, wherein the first region identifier is used to reflect the region to which the user belongs when the audio perception request is initiated. Based on the first region identifier and the first time information, the background collection indicator data corresponding to the audio and video perception request is determined, wherein the background collection indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. Based on the background collected indicator data and the pre-fitted target perception evaluation model, the audio and video perception result corresponding to the audio and video perception request is determined. The target perception evaluation model is used to reflect perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference-plus-noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
[0006] Secondly, embodiments of the present invention also provide a ViNR user audio and video perception assessment device, the device comprising: The first region identifier acquisition module is used to, in response to receiving an audio and video perception request, acquire first time information and a first region identifier from the audio and video perception request, wherein the first region identifier is used to reflect the region to which the user belongs when the audio perception request is initiated. The background collection indicator data determination module is used to determine the background collection indicator data corresponding to the audio and video perception request based on the first area identifier and the first time information. The background collection indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. The audio and video perception result determination module is used to determine the audio and video perception result corresponding to the audio and video perception request based on the background collected index data and the pre-fitted target perception evaluation model; wherein, the target perception evaluation model is used to reflect the perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference-plus-noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
[0007] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the ViNR user audio and video perception evaluation method as described in any embodiment of the present invention.
[0008] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the ViNR user audio and video perception evaluation method as described in any of the embodiments of the present invention.
[0009] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the ViNR user audio and video perception evaluation method as described in any embodiment of the present invention.
[0010] The technical solution of this disclosure firstly, in response to receiving an audio / video perception request, acquires first time information and a first region identifier from the audio / video perception request. Then, based on the first region identifier and the first time information, determines the background acquisition indicator data corresponding to the audio / video perception request. The background acquisition indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. Finally, based on the background acquisition indicator data and a pre-fitted target perception evaluation model, determines the audio / video perception result corresponding to the audio / video perception request. This solves the problem in the prior art where, when determining the audio / video perception result, saving network data from both the sending and receiving ends and using similarity evaluation methods to evaluate various data from both ends, the determination of the audio / video perception result relies on data alignment and synchronization between the two ends. If clock drift, packet loss, or retransmission occurs, comparison errors can easily occur, leading to problems with evaluation accuracy and stability. This disclosure embodiment uses background data collected from the background, such as measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate, to determine the audio and video perception results. This enables rapid location and retrospective analysis of perception problems in specific areas and time periods, improves troubleshooting and optimization efficiency, and achieves a comprehensive, objective, and accurate reflection of the true audio and video perception level of ViNR users. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of exemplary embodiments of the present invention, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the drawings of the embodiments to be described in this invention, and not all of the drawings. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.
[0012] Figure 1 This is a flowchart illustrating a ViNR user audio and video perception evaluation method provided in an embodiment of this disclosure; Figure 2 This is a flowchart illustrating a ViNR user audio and video perception evaluation method provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of a test log recording method provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram illustrating the fitting of the first evaluation result provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram illustrating the fitting of the second evaluation result provided in an embodiment of this disclosure; Figure 6 This is a schematic diagram of the evaluation result adjustment attribute provided in the embodiments of this disclosure; Figure 7A schematic diagram of the structure of a ViNR user audio and video perception evaluation device provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0013] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0014] Before introducing the technical solutions provided by the embodiments of this disclosure, the application scenarios can be illustrated first. The technical solutions provided by the embodiments of this disclosure can be applied to scenarios for evaluating the audio and video perception of ViNR users. Based on the technical solutions of the embodiments of this disclosure, audio and video perception results are determined by collecting background index data such as measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. This enables rapid location and retrospective analysis of perception problems in specific areas and time periods, improves troubleshooting and optimization efficiency, and achieves the effect of comprehensively, objectively, and accurately reflecting the true audio and video perception level of ViNR users.
[0015] Example 1 Figure 1 This is a flowchart illustrating a ViNR user audio and video perception evaluation method provided in this embodiment. This embodiment is applicable to situations where ViNR user audio and video perception is evaluated. The method can be executed by a ViNR user audio and video perception evaluation device, which can be implemented in the form of software and / or hardware. The hardware can be a mobile electronic device, which can execute the ViNR user audio and video perception evaluation method provided in this technical solution.
[0016] like Figure 1 As shown, the method includes: S110. In response to receiving the audio and video perception request, obtain the first time information and the first area identifier in the audio and video perception request.
[0017] The first region identifier is used to reflect the region to which the user belongs when initiating the audio perception request.
[0018] It should be noted that ViNR refers to video calling services based on 5G networks. A ViNR user refers to an end user using ViNR for audio and video calls. An audio / video perception request refers to an evaluation request triggered by the system to assess the audio and video experience of a ViNR user at a given time and place. The user's perception is mainly reflected in their subjectively perceptible experience during the audio / video call. For example, subjectively perceptible aspects of the audio / video call include clarity, smoothness, stuttering, blurriness, audio-visual asynchrony, and dropped calls.
[0019] It should also be noted that "first-time information" refers to the timestamp / time interval information carried in the audio / video perception request. This first-time information is used to pinpoint the specific moment or time period for which the user's perceived experience is being evaluated. "First-region identifier" refers to the identifier indicating the region or location carried in the audio / video perception request. The first-region identifier represents the region where the user was located when the audio perception request was initiated. The first-region identifier typically corresponds to the spatial granularity available for network operations and maintenance. Common optional forms include: cell ID, grid, administrative region, or scene region, etc.
[0020] Specifically, upon receiving an audio / video perception request, the system parses and retrieves the first time information and the first region identifier from the request. The first time information identifies the time point or time period corresponding to the audio / video to be evaluated, so that the corresponding data can be retrieved and aligned on the same time dimension subsequently. The first region identifier reflects the region to which the user is located when initiating the audio / video perception request.
[0021] S120. Based on the first area identifier and the first time information, determine the background collection indicator data corresponding to the audio and video perception request.
[0022] The background data collection metrics include measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate.
[0023] The measurement data consists of data generated by the network side through periodic measurements and statistical analysis of wireless link quality and coverage. This measurement data can serve as reference signal received power, reflecting the coverage strength of the serving cell and the basic quality of the wireless environment. Uplink signal-to-noise ratio (SNR) is a measure of the ratio of useful signal power to interference plus noise power in the cell's uplink. A higher SNR generally indicates less interference in uplink transmission and better link quality. Packet loss rate (PSR) refers to the proportion of service data packets lost during transmission within a given statistical time window. PSR can be obtained based on RTP packet sequence numbers or network-side packet loss counts. An increased PSR often leads to performance degradation such as stuttering, screen tearing, or audio-visual desynchronization.
[0024] It should be noted that regional user air interface latency refers to the statistical value of the latency generated by the transmission and scheduling of user data over the wireless air interface within a certain area. Increased regional user air interface latency usually exacerbates interaction lag, causes playback buffering, and reduces the subjective experience. Regional uplink user plane rate refers to the statistical value of the actual throughput rate of user service data on the uplink user plane within a certain area. Insufficient regional uplink user plane rate can lead to limited encoding bitrate, degraded image quality, or stuttering.
[0025] Specifically, when determining the background collection indicator data corresponding to the audio / video perception request based on the first area identifier and the first time information, the network management granularity and object scope to be queried are first determined based on the first area identifier. For example, the first area identifier is resolved into NR cell ID, sector ID, or grid ID, and the corresponding base station or cell's background collection table, performance counter, or measurement library is located accordingly. Then, based on the first time information, the query time point or time window is determined. For example, a query is performed within a window of several seconds or minutes centered on the first time information to ensure that the obtained indicators are consistent with or closest to the time when the user initiated the audio / video perception request. Further, under the constraints of the first area identifier and the first time information, measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate, etc., within this range are read from the background. If the background data consists of multiple records, they can be aggregated according to preset rules to form the background collection indicator data corresponding to the request.
[0026] S130. Based on the index data collected in the background and the target perception evaluation model obtained in advance, determine the audio and video perception result corresponding to the audio and video perception request.
[0027] Among them, the target perception evaluation model is used to reflect perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference-plus-noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
[0028] It should be noted that the target perception evaluation model is a mathematical model established through pre-testing, subjective scoring, and data fitting. This model maps data collected from the backend, such as measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate, into user-perceptible audio and video perception results. The audio and video perception result refers to the experience evaluation output calculated for a specific audio and video perception request at a given time and region. This result can be represented by a score.
[0029] It's important to note that the reference signal received power dimension is an evaluation dimension that uses measured data as the core independent variable to characterize the impact of coverage strength on perception. A higher reference signal received power dimension generally indicates better coverage and a more stable user experience. The signal-to-interference-plus-noise ratio dimension is an evaluation dimension that uses the uplink signal-to-noise ratio as the core independent variable to characterize the impact of interference and noise on link quality. A higher signal-to-interference-plus-noise ratio dimension generally indicates less interference and better usable link quality.
[0030] It's also important to clarify that the packet loss rate dimension, with packet loss rate as the core independent variable, characterizes the impact of packet loss on the experience degradation caused by audio / video stuttering, screen tearing, and audio-visual asynchrony. A higher packet loss rate generally indicates a worse experience. The air interface latency dimension, with regional user air interface latency as the core independent variable, characterizes the impact of wireless scheduling and transmission latency on real-time audio / video interaction and playback continuity. A higher air interface latency dimension generally indicates a worse interaction and a greater likelihood of buffering and stuttering. The uplink user plane rate dimension, with the actual uplink user plane throughput rate as the core independent variable, characterizes the impact of uplink bandwidth or throughput on uplink video encoding bitrate, clarity, and stability. A lower uplink user plane rate dimension generally indicates worse image quality or a greater likelihood of stuttering.
[0031] Specifically, the system extracts numerical values that match the target perception evaluation model from the background data collected for the audio / video perception request. For example, it extracts measurement data, uplink signal-to-noise ratio (SNR), packet loss rate, regional user air interface latency, and regional uplink user plane rate. Furthermore, it performs unit conversion, missing value handling, and time window aggregation as necessary. Then, it substitutes the measurement data, uplink SNR, packet loss rate, regional user air interface latency, and regional user air interface latency into the pre-fitted target perception evaluation model to obtain the audio / video perception results corresponding to the audio / video perception request.
[0032] Optionally, the first sub-function of the reference signal received power dimension in the target perception evaluation model is determined by fitting in the following manner: in response to the establishment of communication between the calling terminal and the called terminal, while the calling terminal locks onto the NR area and gradually moves away from the NR area, the call video is recorded and the test log is logged; under the condition that the screen recording log time of the call video and the time information of the test log are aligned, the first evaluation result of each sampling point is obtained; based on the multiple first evaluation results of the same sampling point and the reference signal received power of the sampling point, the coefficient values in the first sub-function are fitted and determined.
[0033] The sampling points are related to time information. The first sub-function is a component function in the target perception evaluation model. The first sub-function obtains its specific function form and coefficients through pre-sampling and fitting, and is used to directly convert the reference signal received power collected in the background into a first evaluation result that can be used to evaluate audio and video perception. The calling terminal is the user equipment that initiates the ViNR audio and video call. The calling terminal is used to actively dial, establish a session, and move as required during the test to change wireless environment parameters. The called terminal is the user equipment that receives the ViNR audio and video call. During the test, the called terminal is usually fixed in a location with good wireless conditions to stabilize the other end factors and reduce interference variables. NR area refers to a specific geographical area covered and provided by a 5G wireless network. Call video refers to the visualization result of the actual screen content and playback status presented by the terminal during a ViNR call. Call video can be saved as a video file for subsequent subjective evaluation through screen recording or other methods. Test logs refer to network and service-related data logs recorded by the test software or terminal during the call, typically including timestamps, the reference signal received power at that time, and wireless indicators such as signal-to-interference-plus-noise ratio. Screen recording log time refers to the time information recorded by the screen recording tool for the call video, used to correlate the experiential phenomena in the call video with the specific time points when network metrics occurred. Sampling points can be a specific time point or a time window. Each sampling point corresponds to: the reference signal received power within that time point / time window, and the video segment within that time point / time window. The first evaluation result refers to the perceptual score obtained under the reference signal received power dimension. The first evaluation result can be: the result obtained after scoring the call video according to a predetermined scoring table and mapping it to a unified vMOS scoring system.
[0034] It's important to note that, firstly, the calling and called terminals must initiate and establish a ViNR audio / video call to ensure a stable service link. The called terminal should be placed in a location with excellent signal strength and kept as stationary as possible to minimize the impact of link fluctuations on the user experience. For example, the called terminal can be placed in a location with a reference signal received power > -70 dBm and a signal-to-interference-plus-noise ratio > 25 dB. Testing software should be used to lock the calling terminal onto a designated NR area to prevent sudden changes in performance during handovers, thus making changes in reference signal received power due to distance more controllable. Then, the calling terminal should be gradually moved away from this NR area in an open environment, allowing the reference signal received power to change as continuously and approximately linearly as possible over a range, thus creating multi-level sampling coverage from best to worst.
[0035] During the process of the calling terminal locking onto the NR area and gradually moving away from it, the calling terminal initiates screen recording to capture the entire call video. This ensures consistent screen recording parameters for each test, avoiding user experience differences introduced by the screen recording settings themselves. Simultaneously, test logs are recorded using testing tools, including at least the timestamp and corresponding reference signal reception power for each sampling record, ensuring a stable log time base. After obtaining the call video and test logs, the screen recording log time is aligned with the test log time. Specifically, assuming both the screen recording and testing software use the same system clock, they are directly aligned based on the timestamp. If there is a discrepancy between the screen recording and testing software, a synchronization event can be artificially created during the call, and the time offset can be calibrated using the position where this synchronization event appears in both the call video and the test log. Screen recording segments are extracted based on sampling points, and one or more evaluators score each segment according to a preset scoring table to obtain the first evaluation result for that sampling point.
[0036] It should also be noted that when fitting and determining the coefficient values in the first sub-function, the reference signal received power at that sampling point is used as the independent variable, and the first evaluation result after aggregation at that sampling point is used as the dependent variable. Based on a pre-defined form of the first sub-function, such as a quadratic function, the coefficient values of the first sub-function are fitted using regression methods such as least squares. After fitting, the first sub-function becomes a computable function in the target perception evaluation model, where inputting a reference signal received power will output a corresponding evaluation result, which can be directly used in subsequent online evaluations.
[0037] In this embodiment, the second sub-function of the signal-to-interference-plus-noise ratio dimension in the target perception evaluation model is determined by fitting in the following manner: in response to the establishment of communication between the calling terminal and the called terminal, during the process of the active terminal locking the NR area and moving towards the interference area, the call video is recorded and the test log is recorded; based on the time information of the call video and the recorded test log, the second evaluation result of each sampling point is obtained; based on the multiple second evaluation results of the same sampling point and the signal-to-interference-plus-noise ratio of the sampling point, the coefficient value in the second sub-function is fitted and determined.
[0038] The sampling points are related to time information. The second sub-function is a component function in the target perception evaluation model. The second sub-function, obtained through pre-sampling and fitting, has a specific functional form and coefficients, used to directly convert the signal-to-interference-plus-noise ratio (SNR) collected in the background into a second evaluation result that can be used to evaluate audio and video perception. The interference area refers to a geographical area within the wireless network coverage where, due to factors such as signals from other cells on the same or adjacent frequencies as the target NR area, external electromagnetic sources, or multipath superposition, the interference power received by the user terminal significantly increases, resulting in a significant decrease in the SNR. In the test of this embodiment, the interference area is typically characterized by proximity to co-frequency interfering cells, strong overlapping coverage, or areas near cell boundaries. When the calling terminal moves towards this interference area, it will primarily lower the SNR while keeping the change in the reference signal received power relatively controllable, in order to fit the second sub-function in the SNR dimension. Time information is used to correlate the experiential phenomena in the call video with the specific time points when network indicators occur. The second evaluation result refers to the perception score obtained in the SNR dimension. The second evaluation result can be the result obtained by scoring the video call according to the established scoring table and mapping it to the unified vMOS scoring system.
[0039] It's important to note that, firstly, the calling and called terminals must initiate and establish a ViNR audio / video call to ensure a stable service link. The called terminal should be placed in a location with excellent signal strength and kept as stationary as possible to minimize the impact of link fluctuations on the user experience. Testing software should be used to lock the calling terminal onto a designated NR area to prevent sudden changes in metrics during handover. Then, the calling terminal should be moved towards an area with interference. During this movement, the calling terminal should start screen recording to capture the entire call. Simultaneously, a test log should be recorded using testing tools, including at least the timestamp of each sample and the corresponding signal-to-interference-plus-noise ratio. After obtaining the call video and test log, the timestamps in the screen recording log and the test log should be aligned.
[0040] It should also be noted that when determining the coefficients of the second sub-function, the signal-to-interference-plus-noise ratio at the sampling point is used as the independent variable, and the aggregated second evaluation result at that sampling point is used as the dependent variable. Based on a pre-defined form of the second sub-function, such as a quadratic function, the coefficients of the second sub-function are fitted using regression methods such as least squares. After fitting, the second sub-function becomes a computable function in the target perception evaluation model, which, when given a signal-to-interference-plus-noise ratio, outputs a corresponding perception score for direct use in subsequent online evaluations.
[0041] Optionally, the first and second evaluation results are determined based on the call video quality evaluation results of multiple users at the sampling points.
[0042] It's important to note that for each sampling point, it's not a matter of just one person scoring it; rather, multiple users or reviewers rate the quality of the corresponding call video segment. These ratings are then statistically aggregated to obtain the final first and second evaluation results for that sampling point. Using multiple users' evaluations of the call video quality at each sampling point reduces the error of individual subjective judgment, making the ratings for each sampling point more stable and closer to the overall experience of real users.
[0043] Optionally, the third sub-function of the packet loss rate dimension in the target perception evaluation model is determined in the following way: the call video and recorded test logs are divided according to the preset time slice, and the instantaneous packet loss rate is calculated; the evaluation result adjustment attributes corresponding to different loss rates are determined; and the piecewise function in the third sub-function is created according to the different evaluation result adjustment attributes.
[0044] It should be noted that the third sub-function is a sub-model in the target perception evaluation model used to characterize the corrective effect of packet loss rate on the user's audio and video perception score. The third sub-function typically takes the packet loss rate as input and performs piecewise attenuation or adjustment on the obtained evaluation results. The call video and corresponding test logs are segmented according to preset time slices to calculate the instantaneous packet loss rate within each time slice. The instantaneous packet loss rate refers to the proportion of packet loss calculated within a single time slice window. The instantaneous packet loss rate is usually calculated as the ratio of the number of missing RTP packet sequence numbers to the number that should arrive within that time slice window, reflecting sudden packet loss within a short period. The evaluation result adjustment attribute refers to the descriptive parameters and corresponding weighting coefficients describing the adjustment method and intensity to be applied to the perception score when the instantaneous packet loss rate falls into a certain interval. A piecewise function is a function that divides the range of independent variables into multiple intervals and uses different calculation formulas or different weighting coefficients in different intervals. In this embodiment of the invention, the statistical instantaneous packet loss rate is piecewise correlated with perception; therefore, a piecewise function is used to express different attenuation rules corresponding to different packet loss intervals.
[0045] Specifically, when dividing call videos and recorded test logs according to preset time slices and calculating the instantaneous packet loss rate: First, determine the length of the preset time slice, that is, first set the preset time slice duration. And determine the start and end time references for slicing. Press the recorded video call button. …The video is divided into continuous segments. Each time slice corresponds to a sampling point, which is later used for subjective scoring. Then, the recorded test logs are divided into the same time windows to ensure that log data with a consistent time range can be found for each video segment. Finally, the instantaneous packet loss rate is calculated within each slice; that is, within each time slice, the number of lost packets is counted based on the continuity of the RTP packet sequence number. The instantaneous packet loss rate can be calculated as: Instantaneous Packet Loss Rate = Number of Lost Packets / Theoretically Expected Number of Received Packets. When adjusting attributes to determine the evaluation results corresponding to different loss rates: First, the instantaneous packet loss rate is divided. For example, the instantaneous packet loss rate can be… Divided into: <1%, 1%≤ <3%, 3%≤ <10%, 10%≤ <30% and The algorithm identifies intervals such as ≥30%. Then, for each packet loss rate interval, it combines the perceived scores of the video segments within that interval to determine the appropriate adjustment strength. The adjustment strength is expressed using computable attributes, meaning the evaluation results can be adjusted using a set of weights.
[0046] For example, the piecewise function in the third sub-function can be: ; in, The evaluation results are based on the first and second sub-functions; This refers to the packet loss rate; , , as well as These are the weighting coefficients for the lost intervals of each group; To account for the evaluation results of packet loss fitting, this method is used to correct the ViNR user perception score initially fitted on the wireless side, fully considering the impact of packet loss on ViNR user perception, and the resulting value is closer to the user's perception.
[0047] In this embodiment, the fourth sub-function corresponding to the air interface delay dimension in the target perception evaluation model is based on the following method: determining the first adjustment coefficient of the interval range corresponding to the air interface delay of the NR area; and determining the fourth sub-function based on the first adjustment coefficient.
[0048] It should be noted that the fourth sub-function is the mathematical function in the target perception assessment model used to characterize the impact of air interface delay. The fourth sub-function takes the air interface delay of the NR cell as input and employs different calculation rules for different delay intervals, converting the data obtained in the previous stage... Adjustments are made to obtain the evaluation result after taking time delay into account. The first adjustment coefficient is selected based on the range within which the air interface time delay falls.
[0049] For example, the formula for the fourth sub-function can be: ; in, The evaluation results take into account air interface latency; For NR cell users, the air interface latency is in milliseconds. and This represents the weighting coefficient for the air interface delay interval of each NR cell user.
[0050] In this embodiment, the fifth sub-function corresponding to the uplink user plane rate dimension in the target perception evaluation model is determined in the following way: based on the interval range corresponding to the uplink user plane rate within the region, a second adjustment coefficient for the interval range is determined; and based on the second adjustment coefficient, the fifth sub-function is determined.
[0051] It should be noted that the fifth sub-function is the mathematical function used to characterize the uplink user plane rate in the target perception evaluation model. The fifth sub-function takes the uplink user plane rate as input and uses different calculation rules for different intervals corresponding to different uplink user plane rates, taking the value obtained in the previous stage as input. Adjustments are made to obtain the evaluation result after considering the uplink user plane rate. The second adjustment coefficient is selected based on the range within which the uplink user plane rate falls.
[0052] For example, the formula for the fifth sub-function can be: ; in, This represents the final evaluation result obtained from the fitting process. This refers to the uplink user plane rate, in milliseconds. and These are the weighting coefficients for each uplink user plane rate. The radio-side measurement data and uplink signal-to-noise ratio are initially fitted to obtain preliminary fitting results. These preliminary fitting results are then corrected based on packet loss rate, regional user air interface latency, and regional uplink user plane rate, ultimately yielding an evaluation result that more closely reflects the user's audio and video perception.
[0053] The technical solution of this disclosure firstly, in response to receiving an audio / video perception request, acquires first time information and a first region identifier from the audio / video perception request. Then, based on the first region identifier and the first time information, determines the background acquisition indicator data corresponding to the audio / video perception request. The background acquisition indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. Finally, based on the background acquisition indicator data and a pre-fitted target perception evaluation model, determines the audio / video perception result corresponding to the audio / video perception request. This solves the problem in the prior art where, when determining the audio / video perception result, saving network data from both the sending and receiving ends and using similarity evaluation methods to evaluate various data from both ends, the determination of the audio / video perception result relies on data alignment and synchronization between the two ends. If clock drift, packet loss, or retransmission occurs, comparison errors can easily occur, leading to problems with evaluation accuracy and stability. This disclosure embodiment uses background data collected from the background, such as measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate, to determine the audio and video perception results. This enables rapid location and retrospective analysis of perception problems in specific areas and time periods, improves troubleshooting and optimization efficiency, and achieves a comprehensive, objective, and accurate reflection of the true audio and video perception level of ViNR users.
[0054] Example 2 Figure 2 This is a flowchart illustrating the ViNR user audio and video perception evaluation method provided in this embodiment of the invention. Based on the aforementioned embodiments, it provides a detailed explanation of how to determine the audio perception result corresponding to the audio and video perception request based on the background-collected index data and the pre-fitted target perception evaluation model. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0055] like Figure 2 As shown, the method specifically includes the following steps: S210. In response to receiving the audio and video perception request, obtain the first time information and the first area identifier in the audio and video perception request.
[0056] S220. Based on the first area identifier and the first time information, determine the background collection indicator data corresponding to the audio and video perception request.
[0057] S230. Based on the measurement data in the background indicator data and the first sub-function, determine the first attribute information, and based on the second sub-function and the uplink signal-to-noise ratio, determine the second attribute information.
[0058] The first attribute information is the attribute value obtained by substituting the received power of the reference signal into the first sub-function of the target perception evaluation model, which characterizes the impact of coverage strength on audio and video experience. The second attribute information is the attribute value obtained by substituting the uplink signal-to-noise ratio into the second sub-function of the target perception evaluation model, which characterizes the impact of interference resistance on audio and video experience.
[0059] Specifically, the reference signal received power is retrieved from the background indicator data and used as the input variable of the first sub-function. The first sub-function outputs an experience contribution value corresponding to the reference signal received power, which is the first attribute information. The uplink signal-to-noise ratio is retrieved from the background indicator data and used as the input variable of the second sub-function. The second sub-function outputs an experience contribution value corresponding to the uplink signal-to-noise ratio, which is the second attribute information.
[0060] S240. Determine the third attribute information based on the first attribute information and the second attribute information.
[0061] The third attribute information is the attribute value obtained by processing the first attribute information and the second attribute information through processes such as weighting, combination mapping, or rule synthesis.
[0062] Specifically, the first attribute information and the second attribute information are processed and then fused to obtain the third attribute information. This fusion can be a weighted sum, taking the smaller value as the bottleneck, or generating a comprehensive score by fitting a mapping function, depending on the setting of the coupling relationship between the two in the target perception assessment model.
[0063] S250. Determine the fourth attribute information based on the third attribute information and the target piecewise function corresponding to the group loss rate.
[0064] The fourth attribute information is obtained by combining the third attribute information with the packet loss rate and adjusting it using a target piecewise function corresponding to the packet loss rate.
[0065] Specifically, the target piecewise function takes the packet loss rate as input and outputs an evaluation result adjustment attribute; by applying this adjustment attribute to the third attribute information, the corrected experience value under the current packet loss conditions can be obtained, which is the fourth attribute information.
[0066] S260. Based on the air interface delay of the regional user, determine the first target adjustment coefficient in the fourth sub-function, and determine the fifth attribute information based on the first target adjustment coefficient and the fourth attribute information.
[0067] The fifth attribute information is the attribute value obtained by determining the first target adjustment coefficient in the fourth sub-function based on the air interface delay of the regional user, and then using this coefficient to correct the delay dimension of the fourth attribute information.
[0068] Specifically, the air interface latency of regional users is first determined to fall into a certain latency interval defined by the fourth sub-function. Within this interval, the corresponding first target adjustment coefficient is selected. Then, this coefficient is used to correct the fourth attribute information to obtain the experience value that reflects the impact of latency, which is the fifth attribute information.
[0069] S270. Based on the regional uplink user plane rate, determine the second target adjustment coefficient in the fifth sub-function, and based on the second target adjustment coefficient and the fifth attribute information, determine the sixth attribute information.
[0070] The sixth attribute information is the attribute value obtained by determining the second target adjustment coefficient in the fifth sub-function based on the regional uplink user plane rate, and then using this coefficient to correct the uplink rate dimension of the fifth attribute information.
[0071] Specifically, the uplink user plane rate of the region is first determined to fall into a certain rate range defined by the fifth sub-function, and the corresponding second target adjustment coefficient is selected within this range. Then, this coefficient is used to correct the fifth attribute information to obtain the final experience value after reflecting the impact of uplink carrying capacity, which is the sixth attribute information.
[0072] S280. Determine the audio and video perception results based on the quality assessment results corresponding to the sixth attribute information.
[0073] Specifically, the sixth attribute information is matched with a preset quality rating or mapping rule to obtain the final audio and video perception result.
[0074] The technical solution of this disclosure embodiment, in response to receiving an audio / video perception request, acquires first time information and a first region identifier from the audio / video perception request. Based on the first region identifier and the first time information, determines the background acquisition indicator data corresponding to the audio / video perception request. Then, based on the measurement data and a first sub-function in the background indicator data, determines first attribute information, and based on a second sub-function and uplink signal-to-noise ratio, determines second attribute information. Based on the first and second attribute information, determines third attribute information. Further, based on the third attribute information and the target piecewise function corresponding to the packet loss rate, determines fourth attribute information. Based on the regional user air interface delay, determines a first target adjustment coefficient in the fourth sub-function, and based on the first target adjustment coefficient and the fourth attribute information, determines fifth attribute information. Based on the regional uplink user plane rate, determines a second target adjustment coefficient in the fifth sub-function, and based on the second target adjustment coefficient and the fifth attribute information, determines a sixth attribute information. Finally, based on the quality assessment result corresponding to the sixth attribute information, determines the audio / video perception result. By incorporating multi-dimensional network key indicators such as measurement data, uplink signal-to-noise ratio, packet loss rate, air interface latency, and uplink user plane rate into a unified target-aware evaluation model, and calculating attribute information step by step from basic wireless quality to packet loss correction to latency correction and then to rate correction, the evaluation process can comprehensively characterize the combined impact of coverage, interference, transmission reliability, and carrying capacity on audio and video experience, thereby improving the accuracy and interpretability of the evaluation results.
[0075] Example 3 As an optional embodiment of the present invention, an example is provided to further illustrate the invention.
[0076] It should be noted that the main indicators reflecting the quality of the wireless environment in wireless networks are the reference signal received power (RSW) and uplink signal-to-noise ratio (SNR). The ViNR video sensing scheme uses RSW and SNR as the basic indicators for preliminary fitting. By analyzing the influence of RSW and SNR on vMOS, a preliminary fitting relationship is obtained, and then the results are corrected through analysis of other factors. For the influence of a single wireless indicator on vMOS, this scheme uses the controlled variable method for testing and analysis. For network problems affected by multiple factors, the method of controlling factors is used to transform the multi-factor problem into multiple single-factor problems. The remaining factors are kept as constant as possible, and only one factor is changed at a time to study the impact of the changed factor on the phenomenon, and these are studied separately.
[0077] To analyze the relationship between the received reference signal power and vMOS in wireless communication, it is necessary to reduce or even avoid the impact of external wireless interference. Therefore, a simple indoor cell coverage scenario with a simple wireless environment was selected for testing. The test procedure involved fixing the called terminal at a location with excellent signal strength, while the calling terminal locked onto the NR serving cell. The device was then gradually moved away from the serving cell in an open area to allow the received reference signal power to change linearly. Screen recording software was used to record the ViNR image quality and smoothness, and a test log was also recorded. The screen recording software parameters were kept consistent for each test. Subsequent analysis will use the received reference signal power as a reference axis, aligning the timestamps of the test log and screen recording log to conduct user subjective scoring. For details on the scoring method, please refer to [link to test scoring method]. Figure 3 See also Figure 3 The scoring method employs a no-reference evaluation approach, requiring no original reference video. Users score the video based on extracted distortion effects, using a 10-point scale. Users subjectively score audio and video smoothness features. For example: 10 points for perceived high definition; 9 points for perceived standard definition; 8 points for perceived smoothness; 7 points for perceived glitches or blurriness; 6 points for perceived stuttering; 4 points for perceived audio and video desynchronization; 2 points for perceived video freezes or silence; and 1 point for perceived dropped calls. The scores for each sampling point are obtained through this method and then mapped to a 5-point vMOS system to obtain a preliminary subjective score. .
[0078] The subjective scores were fitted and then mapped to a 5-point scoring system. The fitting mapping relationship is as follows: ; in, The number of sampling points. The score for each sampling point, This is the initial fitted score. When the audio and video are high definition, the corresponding score is... The value is 5; when the audio and video quality degrade is perceptible, but the video is smooth, the corresponding value is 5. The value is 4; when the audio and video are smooth, but there are slight stutters or screen glitches, the corresponding value is 4. The value is 3; when the audio and video are offensive but tolerable, the corresponding value is... The value is 2; when the audio and video are very offensive and unbearable, the corresponding value is 2. The value is 1. This evaluation scheme is simpler to implement than other schemes, more suitable for real-time evaluation of video quality, and can improve the accuracy of the evaluation scheme by increasing the number of test sampling points and reducing subjective judgment errors. Test data integration: Based on the above evaluation scheme, multiple sampled samples are scored, and the scores for different reference signal received powers are integrated. See [link to relevant documentation]. Figure 4 The fitting relationship between the reference signal received power (x) and the vMOS value (y) was obtained.
[0079] When fitting uplink signal-to-noise ratio (SNR) test data for wireless communication, to reduce the impact of reference signal received power on uplink SNR sampling points, outdoor cells with high wireless overlap coverage and significant co-channel interference were selected for testing. The called terminal was placed in a location with excellent signal strength. The test plan involved locking the calling terminal onto the serving cell using test software, and then gradually moving it away from the serving cell in an open area towards a co-channel interfering cell to make the uplink SNR tend to change linearly. Simultaneously, to enrich the sampling points, some data were tested using a vehicle-mounted method without locking onto the serving cell. For the analysis of the impact of sampling and fitting methods on reference signal received power, please refer to [link to relevant documentation]. Figure 5 Finally, the fitting relationship between the uplink signal-to-noise ratio (x) and vMOS (y) is obtained.
[0080] The impact of changes in the wireless environment on call quality during ViNR video calls is obvious. When wireless quality is good, users are unlikely to perceive differences in call quality. However, when wireless quality deteriorates, video quality also declines, directly affecting user experience. During wireless optimization, front-end testing metrics cannot comprehensively assess wireless quality; typically, back-end wireless metrics are used to evaluate the wireless network. Therefore, it is necessary to predict ViNR video user experience by analyzing changes in back-end wireless metrics, providing a valuable basis for subsequent optimization measures and complaint handling.
[0081] In the previous step, the test indicators reference signal received power & uplink signal-to-noise ratio and 5-point subjective scoring were obtained through front-end wireless testing and subjective scoring. The next step in fitting the scores is to correlate the front-end wireless test KPIs with the back-end wireless network metrics, completing a crucial mapping step. Based on the fitting results of the test data from the first step, the front-end test metrics, reference signal received power and uplink signal-to-noise ratio (SNR), are mapped to the back-end wireless network metrics. The back-end wireless network metric corresponding to the front-end test reference signal received power is the reference signal received power, defined as the linear average power of resource units carrying the cell-specific reference signal in the frequency band being measured, and is a key indicator reflecting the coverage of the serving cell. The back-end wireless network metric corresponding to the front-end test uplink SNR is the uplink SNR, defined as the uplink SNR for all users in the cell. The back-end wireless network metrics need to be fitted according to the fitting relationship of the front-end test metrics to obtain the ViNR user perception score fitted by the back-end wireless network metrics. ;in, The ViNR user perception score is fitted to the background wireless network metrics. Factors affecting wireless network KPIs , as well as These are the network relationship fitting coefficients. The fitting relationship from the front-end wireless test is mapped to the wireless network KPIs, and the reference signal received power and uplink signal-to-noise ratio are initially fitted to the ViNR user perception score. .
[0082] Preliminary fitting The current value only considers the impact of two key KPIs of wireless networks: reference signal received power and uplink signal-to-noise ratio. However, in actual user video perception, in addition to the influence of network signal strength, the impact of key factors such as packet loss rate, regional user air interface latency, and regional uplink user plane rate also needs to be considered. To obtain a score that more closely approximates the user's actual perception, it is necessary to utilize the impact of factors such as packet loss rate, regional user air interface latency, and regional uplink user plane rate on video user perception. The value is corrected to obtain a vMOS value that is closer to the user's actual perception.
[0083] Impact of Packet Loss Rate on User Perception: By observing a large number of video calls and recording the time points of stuttering, the packet loss situation at these stuttering points was analyzed, and the correlation between instantaneous packet loss and video perception was established. Analysis Plan: The test video and recorded logs were segmented into 5-second segments. The instantaneous packet loss rate was calculated using RTP packet sequence numbers. Then, a user subjective scoring system was used to complete the subjective scoring of the video call under different packet loss rates, such as... Figure 6As shown. Due to the large statistical error in instantaneous packet loss rate, the scoring is divided into intervals based on the packet loss rate. The impact of packet loss on video perception is incorporated into the mathematical model. Since the current packet loss rate is piecewise correlated with user perception, a piecewise regression method is used to determine the optimal mathematical model and the polynomial packet loss fitting coefficient of this model, which ultimately determines the piecewise function in the third sub-function.
[0084] By using the above piecewise regression method to perform weight analysis on other factors, the network weight factors of regional user air interface latency and regional uplink user plane rate are also fitted into the mathematical model. The model is then further modified in a positive or negative direction to finally obtain the vMOS value.
[0085] The technical solution of this disclosure embodiment can obtain specific values by grouping and integrating statistical data from test sampling points. The initial fitting of the reference signal received power and uplink signal-to-noise ratio on the wireless side is corrected by adding factors such as packet loss rate, regional user air interface latency, and regional uplink user plane rate, ultimately resulting in an objective reference mathematical model that more closely approximates user video perception, thus improving the accuracy and interpretability of the evaluation results.
[0086] Example 4 Figure 7 This is a schematic diagram of the structure of the ViNR user audio and video perception evaluation device provided in the embodiments of this disclosure, as shown below. Figure 7 As shown, the device includes: a first area identifier acquisition module 310, a background collection indicator data determination module 320, and an audio and video perception result determination module 330.
[0087] A first region identifier acquisition module is used to, in response to receiving an audio / video perception request, acquire first time information and a first region identifier from the audio / video perception request, wherein the first region identifier is used to reflect the region to which the user belongs when the audio / video perception request is initiated; a background collection indicator data determination module is used to determine the background collection indicator data corresponding to the audio / video perception request based on the first region identifier and the first time information, wherein the background collection indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface delay, and regional uplink user plane rate; an audio / video perception result determination module is used to determine the audio / video perception result corresponding to the audio / video perception request based on the background collection indicator data and a pre-fitted target perception evaluation model; wherein the target perception evaluation model is used to reflect perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference plus noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
[0088] The technical solution of this disclosure firstly, in response to receiving an audio / video perception request, acquires first time information and a first region identifier from the audio / video perception request. Then, based on the first region identifier and the first time information, determines the background acquisition indicator data corresponding to the audio / video perception request. The background acquisition indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. Finally, based on the background acquisition indicator data and a pre-fitted target perception evaluation model, determines the audio / video perception result corresponding to the audio / video perception request. This solves the problem in the prior art where, when determining the audio / video perception result, saving network data from both the sending and receiving ends and using similarity evaluation methods to evaluate various data from both ends, the determination of the audio / video perception result relies on data alignment and synchronization between the two ends. If clock drift, packet loss, or retransmission occurs, comparison errors can easily occur, leading to problems with evaluation accuracy and stability. This disclosure embodiment uses background data collected from the background, such as measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate, to determine the audio and video perception results. This enables rapid location and retrospective analysis of perception problems in specific areas and time periods, improves troubleshooting and optimization efficiency, and achieves a comprehensive, objective, and accurate reflection of the true audio and video perception level of ViNR users.
[0089] Based on the above technical solutions, the first sub-function of the reference signal received power dimension in the target perception evaluation model is determined by fitting in the following manner: In response to the establishment of communication between the calling terminal and the called terminal, while the calling terminal locks onto the NR area and gradually moves away from the NR area, a call video is recorded and a test log is recorded; under the condition that the screen recording log time of the call video is aligned with the time information of the recorded test log, the first evaluation result of each sampling point is obtained; wherein, the sampling point is related to the time information; based on multiple first evaluation results of the same sampling point and the reference signal received power of the sampling point, the coefficient values in the first sub-function are fitted and determined.
[0090] Based on the above technical solutions, the second sub-function of the signal-to-interference-plus-noise ratio dimension in the target perception evaluation model is determined by fitting in the following manner: In response to the establishment of communication between the calling terminal and the called terminal, during the process of the active terminal locking the NR area and moving towards the interference area, the call video is recorded and the test log is recorded; according to the time information of the call video and the recorded test log, the second evaluation result of each sampling point is obtained, wherein the sampling point is related to the time information; based on multiple second evaluation results of the same sampling point and the signal-to-interference-plus-noise ratio of the sampling point, the coefficient value in the second sub-function is fitted and determined.
[0091] Based on the above technical solutions, the first evaluation result and the second evaluation result are determined based on the call video quality evaluation results of multiple users at the sampling points.
[0092] Based on the above technical solutions, the third sub-function of the packet loss rate dimension in the target perception evaluation model is determined in the following way: the call video and recorded test log are divided according to a preset time slice, and the instantaneous packet loss rate is calculated; the evaluation result adjustment attribute corresponding to different loss rates is determined; and the segmentation function in the third sub-function is created according to the different evaluation result adjustment attributes.
[0093] Based on the above technical solutions, the fourth sub-function corresponding to the air interface delay dimension in the target perception evaluation model is based on the following method: determining the first adjustment coefficient of the interval range corresponding to the air interface delay of the NR area; and determining the fourth sub-function based on the first adjustment coefficient.
[0094] Based on the above technical solutions, the fifth sub-function corresponding to the uplink user plane rate dimension in the target perception evaluation model is determined in the following way: based on the interval range corresponding to the uplink user plane rate within the region, a second adjustment coefficient of the interval range is determined; and the fifth sub-function is determined according to the second adjustment coefficient.
[0095] Based on the above technical solutions, the audio and video perception result determination module 330 is further configured to determine first attribute information based on the measurement data and the first sub-function in the background indicator data, and determine second attribute information based on the second sub-function and the uplink signal-to-noise ratio; determine third attribute information based on the first attribute information and the second attribute information; determine fourth attribute information based on the third attribute information and the target piecewise function corresponding to the packet loss rate; determine a first target adjustment coefficient in the fourth sub-function based on the regional user air interface delay, and determine fifth attribute information based on the first target adjustment coefficient and the fourth attribute information; determine a second target adjustment coefficient in the fifth sub-function based on the regional uplink user surface rate, and determine a sixth attribute information based on the second target adjustment coefficient and the fifth attribute information; and determine the audio and video perception result based on the quality assessment result corresponding to the sixth attribute information.
[0096] The ViNR user audio and video perception assessment device provided in this disclosure can execute the ViNR user audio and video perception assessment method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.
[0097] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0098] Example 5 Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Refer to the following... Figure 8 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 8 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals). Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0099] like Figure 8 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0100] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0101] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0102] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0103] The electronic device provided in this embodiment and the ViNR user audio and video perception evaluation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0104] Example 6 This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the ViNR user audio and video perception evaluation method provided in the above embodiments.
[0105] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0106] In some implementations, the server may communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and may interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0107] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0108] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: In response to receiving an audio / video perception request, the system obtains first time information and a first region identifier from the audio / video perception request, wherein the first region identifier is used to reflect the region to which the user belongs when the audio perception request is initiated. Based on the first region identifier and the first time information, the background collection indicator data corresponding to the audio and video perception request is determined, wherein the background collection indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. Based on the background collected indicator data and the pre-fitted target perception evaluation model, the audio and video perception result corresponding to the audio and video perception request is determined. The target perception evaluation model is used to reflect perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference-plus-noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
[0109] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0111] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0112] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0113] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0114] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0115] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0116] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A ViNR user audio and video perception assessment method, characterized in that, include: In response to receiving an audio / video perception request, the system obtains first time information and a first region identifier from the audio / video perception request, wherein the first region identifier is used to reflect the region to which the user belongs when the audio perception request is initiated. Based on the first region identifier and the first time information, the background collection indicator data corresponding to the audio and video perception request is determined, wherein the background collection indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. Based on the background collected indicator data and the pre-fitted target perception evaluation model, the audio and video perception result corresponding to the audio and video perception request is determined. The target perception evaluation model is used to reflect perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference-plus-noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
2. The method according to claim 1, characterized in that, The first sub-function of the reference signal received power dimension in the target perception evaluation model is determined by fitting in the following manner: In response to the establishment of communication between the calling terminal and the called terminal, during the process of locking the NR area based on the calling terminal and gradually moving away from the NR area, the call video is recorded and the test log is recorded. Under the condition that the screen recording log time of the call video is aligned with the time information of the recorded test log, the first evaluation result of each sampling point is obtained; wherein, the sampling point is related to the time information; Based on multiple first evaluation results at the same sampling point and the reference signal received power at the sampling point, the coefficient values in the first sub-function are determined by fitting.
3. The method according to claim 1, characterized in that, The second sub-function of the signal-to-interference-plus-noise ratio dimension in the target perception evaluation model is determined by fitting in the following manner: In response to the establishment of communication between the calling terminal and the called terminal, during the process of locking the NR area based on the active terminal and moving towards the interference area, the call video is recorded and the test log is logged. Based on the time information of the call video and the recorded test log, a second evaluation result is obtained for each sampling point, wherein the sampling point is related to the time information; Based on multiple second evaluation results of the same sampling point and the signal-to-interference-plus-noise ratio of the sampling point, the coefficient values in the second sub-function are determined by fitting.
4. The method according to claim 2 or 3, characterized in that, The first and second evaluation results were determined based on the call video quality evaluation results of multiple users at the sampled points.
5. The method according to claim 1, characterized in that, The third sub-function of the packet loss rate dimension in the target perception evaluation model is determined based on the following method: The call video and recorded test logs are divided according to preset time slices, and the instantaneous packet loss rate is calculated. Determine the adjustment attributes for the evaluation results corresponding to different loss rates; Adjust the attributes based on different evaluation results to create a piecewise function in the third sub-function.
6. The method according to claim 1, characterized in that, The fourth sub-function corresponding to the air interface delay dimension in the target perception evaluation model is based on the following method: Based on the range of air interface delay corresponding to the NR area, a first adjustment coefficient for the range is determined; The fourth sub-function is determined based on the first adjustment coefficient.
7. The method according to claim 1, characterized in that, The fifth sub-function corresponding to the uplink user plane rate dimension in the target perception evaluation model is determined based on the following method: A second adjustment coefficient is determined based on the interval range corresponding to the uplink user plane rate within the region. The fifth sub-function is determined based on the second adjustment coefficient.
8. The method according to claim 1, characterized in that, The step of determining the audio perception result corresponding to the audio and video perception request based on the background collected indicator data and the pre-fitted target perception evaluation model includes: Based on the measurement data and the first sub-function in the background indicator data, the first attribute information is determined, and based on the second sub-function and the uplink signal-to-noise ratio, the second attribute information is determined; The third attribute information is determined based on the first attribute information and the second attribute information; The fourth attribute information is determined based on the third attribute information and the target piecewise function corresponding to the packet loss rate; Based on the air interface delay of the user in the region, the first target adjustment coefficient in the fourth sub-function is determined, and the fifth attribute information is determined based on the first target adjustment coefficient and the fourth attribute information. Based on the uplink user plane rate in the region, determine the second target adjustment coefficient in the fifth sub-function, and based on the second target adjustment coefficient and the fifth attribute information, determine the sixth attribute information; The audio and video perception result is determined based on the quality assessment result corresponding to the sixth attribute information.
9. A ViNR user audio and video perception evaluation device, characterized in that, include: The first region identifier acquisition module is used to, in response to receiving an audio and video perception request, acquire first time information and a first region identifier from the audio and video perception request, wherein the first region identifier is used to reflect the region to which the user belongs when the audio perception request is initiated. The background collection indicator data determination module is used to determine the background collection indicator data corresponding to the audio and video perception request based on the first area identifier and the first time information. The background collection indicator data includes measurement data, uplink signal-to-noise ratio, packet loss rate, regional user air interface latency, and regional uplink user plane rate. The audio and video perception result determination module is used to determine the audio and video perception result corresponding to the audio and video perception request based on the background collected index data and the pre-fitted target perception evaluation model; wherein, the target perception evaluation model is used to reflect the perception evaluation information in at least the dimensions of reference signal received power, signal-to-interference-plus-noise ratio, packet loss rate, air interface delay, and uplink user plane rate.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the ViNR user audio and video perception evaluation method as described in any one of claims 1-8.