Large model driven quality experience evaluation method for intelligent agent communication network
By constructing a quality experience evaluation method driven by a large language model in a multi-agent communication network, and utilizing multi-dimensional state feature vectors and data augmentation techniques, the problem of QoE quantification is solved, and accurate and efficient evaluation of user experience quality is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, QoS metrics are difficult to fully cover the quality of user experience (QoE), especially in multi-agent communication networks. There is a complex nonlinear mapping between traditional QoS metrics and user perception, and there is a lack of usable and reliable metrics, which leads to optimization deviating from the user experience.
By obtaining multi-dimensional state feature vectors, an initial labeled dataset is constructed, and the dataset is augmented using a large language model. A service reliability score mapping model is trained, and a service quality assessment service reliability score is output, thus constructing a mapping relationship from system state to user experience quality score.
It enables quantitative evaluation of QoE, improves evaluation efficiency, reduces reliance on manual annotation, lowers data collection costs, and significantly improves prediction accuracy and the comprehensiveness of evaluation.
Smart Images

Figure CN121858982A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication network service quality assessment technology, and in particular to a large model-driven quality experience assessment method for intelligent agent communication networks. Background Technology
[0002] In recent years, with the successful application of generative artificial intelligence in various fields such as communication networks, drone dispatching, and smart city construction, multi-agent communication networks (ACNs) for heterogeneous intelligent agents are becoming the infrastructure for global interconnection and on-demand empowerment. The rapid development of emerging businesses such as AIGC (Artificial Intelligence Generated Content) is profoundly reshaping the user-centric experience quality requirements, prompting academia to shift from Quality of Service (QoS) to Quality of Experience (QoE), emphasizing the integration of network metrics and user subjective feelings.
[0003] In related technologies, traditional QoS indicators (such as latency and energy consumption) are based on objective and quantifiable network parameters and have a relatively complete standardized measurement system for network performance, which is crucial for real-time interaction.
[0004] However, due to the complex nonlinear mapping between QoS metrics and user perception, it is difficult to fully cover the AIGC experience dimension. In addition, the lack of available and reliable reliability metrics makes it easy for optimization to deviate from the user experience, which urgently needs to be addressed. Summary of the Invention
[0005] This application provides a large model-driven quality experience evaluation method for intelligent agent communication networks to solve the problem of difficulty in quantifying QoE evaluation and improve evaluation efficiency.
[0006] The first aspect of this application provides a method for evaluating the quality experience of large-model-driven intelligent agent communication networks, including the following steps: Obtain multidimensional state feature vectors, and construct an initial labeled dataset based on the multidimensional state feature vectors; The initial labeled dataset is augmented based on a pre-defined large language model to obtain an augmented dataset, and a pre-built service reliability score mapping model is trained using the augmented dataset to obtain a trained service reliability score mapping model. The multidimensional state feature vector is input into the trained service reliability score mapping model, and the service reliability score is output through the trained service reliability score mapping model.
[0007] Optionally, the multidimensional state feature vector and the initial labeled dataset include: Collect state samples of multiple intelligent agent devices, wherein the state samples include the latency of intelligent agent devices, the energy consumption of intelligent agent devices, the bandwidth allocated to intelligent agent devices, the task computation of intelligent agent devices, and the total number of intelligent agent devices; Based on a preset function, the state samples of the multiple intelligent agent devices are processed to obtain the multidimensional state feature vector.
[0008] Optionally, constructing the initial labeled dataset based on the multidimensional state feature vector includes: Multiple baseline scenarios were identified; Based on a preset efficient optimization strategy, the first reliability level state of each benchmark scenario is generated, and based on a preset suboptimal strategy with perturbation, the second reliability level state of each benchmark scenario is generated, and based on a preset completely random strategy, the third reliability level state of each benchmark scenario is generated. Obtain the scoring results generated based on the first reliability level state, the second reliability level state, and the third reliability level state, and obtain the initial labeled dataset based on the scoring results.
[0009] Optionally, the step of augmenting the initial labeled dataset based on a preset large language model to obtain an augmented dataset, and then using the augmented dataset to train a pre-built service reliability score mapping model to obtain a trained service reliability score mapping model, includes: A validation set is determined based on the initial labeled dataset, and a probabilistic reference boundary is constructed based on the validation set; Gaussian noise is injected into the multidimensional state feature vector to obtain unlabeled samples; Based on the preset large language model, the nonlinear scoring rules of the original samples in the initial labeled dataset are summarized to generate a reliability score for the unlabeled samples; The reliability score is evaluated based on the probabilistic reference boundary, the preset reasonable interval, and the validation set. If the quality evaluation result does not meet the preset conditions, the preset large language model is iterated based on the quality evaluation result to obtain the augmented dataset. Based on the augmented dataset, reliability score labels are generated, and the service reliability score mapping model is trained based on the reliability score labels.
[0010] Optionally, the step of performing a quality assessment on the reliability score based on a preset reasonable range includes: The precision of the reliability score is calculated based on the augmented dataset, and the recall of the reliability score is calculated based on the validation set. Calculate the harmonic mean of the reliability score based on the recall rate and the precision rate; Whether to iterate the harmonic mean of the reliability score is determined based on a preset harmonic mean threshold.
[0011] Optionally, the multidimensional state feature vector includes basic performance features, task allocation statistics, bandwidth resource allocation statistics, and system configuration features.
[0012] A second aspect of this application provides a large model-driven quality experience evaluation device for intelligent agent communication networks, comprising: The acquisition module is used to acquire multidimensional state feature vectors and construct an initial labeled dataset based on the multidimensional state feature vectors. The training module is used to perform data augmentation on the initial labeled dataset based on a preset large language model to obtain an augmented dataset, and to use the augmented dataset to train a pre-built service reliability score mapping model to obtain a trained service reliability score mapping model. The evaluation module is used to input multi-dimensional state feature vectors into the trained service reliability score mapping model, and output a service quality assessment service reliability score through the trained service reliability score mapping model.
[0013] Optionally, the multidimensional state feature vector and the initial labeled dataset include: Collect state samples of multiple intelligent agent devices, wherein the state samples include the latency of intelligent agent devices, the energy consumption of intelligent agent devices, the bandwidth allocated to intelligent agent devices, the task computation of intelligent agent devices, and the total number of intelligent agent devices; Based on a preset function, the state samples of the multiple intelligent agent devices are processed to obtain the multidimensional state feature vector.
[0014] Optionally, the acquisition module is specifically used for: Multiple baseline scenarios were identified; Based on a preset efficient optimization strategy, the first reliability level state of each benchmark scenario is generated, and based on a preset suboptimal strategy with perturbation, the second reliability level state of each benchmark scenario is generated, and based on a preset completely random strategy, the third reliability level state of each benchmark scenario is generated. Obtain the scoring results generated based on the first reliability level state, the second reliability level state, and the third reliability level state, and obtain the initial labeled dataset based on the scoring results.
[0015] Optionally, the training module is specifically used for: A validation set is determined based on the initial labeled dataset, and a probabilistic reference boundary is constructed based on the validation set; Gaussian noise is injected into the multidimensional state feature vector to obtain unlabeled samples; Based on the preset large language model, the nonlinear scoring rules of the original samples in the initial labeled dataset are summarized to generate a reliability score for the unlabeled samples; The reliability score is evaluated based on the probabilistic reference boundary, the preset reasonable interval, and the validation set. If the quality evaluation result does not meet the preset conditions, the preset large language model is iterated based on the quality evaluation result to obtain the augmented dataset. Based on the augmented dataset, reliability score labels are generated, and the service reliability score mapping model is trained based on the reliability score labels.
[0016] Optionally, the training module is specifically used for: The precision of the reliability score is calculated based on the augmented dataset, and the recall of the reliability score is calculated based on the validation set. Calculate the harmonic mean of the reliability score based on the recall rate and the precision rate; Whether to iterate the harmonic mean of the reliability score is determined based on a preset harmonic mean threshold.
[0017] Optionally, the multidimensional state feature vector includes basic performance features, task allocation statistics, bandwidth resource allocation statistics, and system configuration features.
[0018] A third aspect of this application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to perform a large model-driven quality experience evaluation method for intelligent agent communication networks as described in the above embodiments.
[0019] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the large model-driven quality experience evaluation method for intelligent agent communication networks as described in the above embodiments.
[0020] A fifth aspect of this application provides a computer program product storing a computer program that, when executed by a processor, implements the large model-driven quality experience evaluation method for intelligent agent communication networks as described in the above embodiments.
[0021] Therefore, this embodiment of the application obtains multi-dimensional state feature vectors and constructs an initial labeled dataset based on these vectors. Data augmentation is then performed on the initial labeled dataset using a pre-defined large language model to obtain an augmented dataset. This augmented dataset is then used to train a pre-constructed service reliability score mapping model, resulting in a trained service reliability score mapping model. The multi-dimensional state feature vectors are input into the trained service reliability score mapping model, which then outputs a service reliability score for quality of service evaluation. Thus, by employing a semi-automatic label expansion mechanism assisted by a large language model, a mapping relationship from system state to user experience quality score is constructed, solving the problem of difficulty in quantifying QoE evaluation and improving evaluation efficiency.
[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0023] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a large model-driven quality experience evaluation method for intelligent agent communication networks provided according to an embodiment of this application; Figure 2 This is a flowchart of an LLM data augmentation method for evaluating quality experience in a large model-driven intelligent agent communication network according to an embodiment of this application; Figure 3 This is a schematic diagram of a large model-driven quality experience evaluation device for intelligent agent communication networks provided according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0024] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0025] Before introducing the large model-driven quality experience evaluation method for agent communication networks in this application, let's briefly introduce the large model-driven quality experience evaluation method for agent communication networks in related technologies.
[0026] Specifically, when the service quality assessment methods in related technologies are applied in ACN networks, they face the following key challenges: The evaluation metrics are multidimensional and exhibit nonlinear mappings: the various influencing factors of QoE are highly coupled, and their interactions often exhibit counterintuitive nonlinear relationships. Therefore, no single factor can effectively predict the overall experience, and their complex synergistic or attritional effects must be considered.
[0027] Training data collection is difficult: QoE metrics rely heavily on subjective annotation by experts, which is costly and difficult to implement on a large scale. In addition, the data itself has many dimensions and comes from diverse sources, which together make the collection work extremely challenging.
[0028] Subjectivity and individual differences: QoE is essentially a subjective experience, but centralized assessment methods ignore individual differences among users (such as preferences, environment, and tolerance), making it difficult to truly reflect the diverse perceptual experiences of each individual.
[0029] Evaluation criteria and modeling deficiencies: Current QoE research lacks a universally applicable evaluation paradigm, and relevant technical indicators are difficult to directly map to user experience. Furthermore, confounding variables such as content semantics and user interests further increase the difficulty of modeling. QoE modeling needs to comprehensively consider the coupling relationships between indicators and their semantic impact, and the applicability of traditional linear modeling methods is significantly reduced in this scenario.
[0030] This application addresses the aforementioned problems by proposing a large-model-driven quality experience evaluation method for intelligent agent communication networks. In this method, embodiments of this application obtain multi-dimensional state feature vectors and construct an initial labeled dataset based on these vectors. The initial labeled dataset is then augmented using a pre-defined large language model to obtain an augmented dataset. This augmented dataset is then used to train a pre-constructed service reliability score mapping model, resulting in a trained service reliability score mapping model. The multi-dimensional state feature vectors are input into the trained service reliability score mapping model, which then outputs a service reliability score as a quality evaluation result. Thus, by employing a semi-automatic label expansion mechanism assisted by a large language model, a mapping relationship from system state to user experience quality score is constructed, solving the problem of difficulty in quantifying QoE evaluation and improving evaluation efficiency.
[0031] Specifically, Figure 1 This is a flowchart illustrating a large-model-driven quality experience evaluation method for intelligent agent communication networks provided in an embodiment of this application.
[0032] like Figure 1 As shown, this large-model-driven quality experience evaluation method for agent communication networks includes the following steps: In step S101, a multidimensional state feature vector is obtained, and an initial labeled dataset is constructed based on the multidimensional state feature vector. Optionally, in some embodiments, an initial labeled dataset is constructed based on a multi-dimensional state feature vector, including: determining multiple benchmark scenarios; generating a first reliability level state for each benchmark scenario based on a preset efficient optimization strategy, generating a second reliability level state for each benchmark scenario based on a preset perturbation-based suboptimal strategy, and generating a third reliability level state for each benchmark scenario based on a preset completely random strategy; obtaining scoring results generated based on the first reliability level state, the second reliability level state, and the third reliability level state, and obtaining the initial labeled dataset based on the scoring results.
[0033] Understandably, multi-dimensional state feature vectors are collected from various agent devices in a multi-agent communication network; a seed dataset of 90 high-quality labeled data is constructed. Ten basic configurations covering low, medium, and high heterogeneity are defined, with three scales of 5, 10, and 30, totaling 30 baseline scenarios. In each scenario, three types of policies are run to form significantly graded system states: efficient optimization, perturbated suboptimal, and completely random, corresponding to high reliability level (first reliability level state), medium reliability level (second reliability level state), and low reliability level (third reliability level state). Specifically, efficient optimization (30 data points) collects the states of each device under real-world application policies as high-reliability samples; perturbated suboptimal (30 data points) applies moderate-intensity Gaussian noise to the high-reliability policy to simulate an acceptable suboptimal state; and completely random (30 data points) uses completely random assignment to simulate unreliable states caused by deteriorating conditions or errors. Subsequently, experts subjectively score the 90 states to form the initial labeled dataset. ,in This represents the reliability score.
[0034] Optionally, in some embodiments, the multidimensional state feature vector and the initial labeled dataset include: collecting state samples of multiple intelligent agent devices, wherein the state samples include the latency of the intelligent agent devices, the energy consumption of the intelligent agent devices, the bandwidth allocated to the intelligent agent devices, the computational load of the intelligent agent devices, and the total number of intelligent agent devices; and processing the state samples of the multiple intelligent agent devices based on a preset function to obtain a multidimensional state feature vector.
[0035] Among them, latency is the delay time of data packet transmission; energy consumption is the energy consumed by the intelligent agent device to perform the task; allocated bandwidth is the data transmission rate allocated by the network to the intelligent agent device; task computation is the scale or complexity of the computational task that the intelligent agent device needs to process; and total number of devices is the total number of intelligent agent devices participating in the cooperation in the current network.
[0036] Understandably, state samples are obtained from multiple intelligent agent devices. The parameters of these devices are objective and quantifiable service quality indicators, forming the raw data foundation for evaluating Quality of Experience (QoE). This data is then processed based on a preset function. The system state feature vector (multidimensional state feature vector) has a total of 20 dimensions; among them, Indicates system latency. Indicates system energy consumption. This indicates the bandwidth allocated to the intelligent agent device. I represents the computational workload of the device task, and I represents the total number of intelligent agent devices.
[0037] Therefore, this invention proposes an innovative framework that integrates semantic understanding and nonlinear modeling. It uses LLM to parse context and generate personalized perceptual representations to capture subjective preference differences. By combining SVR and kernel functions to establish a complex mapping between multi-dimensional influencing factors and user experience, it achieves accurate quantification of individual differences in perception, thereby focusing on subjectivity and individual differences.
[0038] Optionally, in some embodiments, the multidimensional state feature vector includes basic performance features, task allocation statistics, bandwidth resource allocation statistics, and system configuration features.
[0039] Understandably, the 20-dimensional numerical feature vector (multidimensional state feature vector) for time-series performance data consists of fundamental performance features, task allocation statistical features, bandwidth resource allocation statistical features, and system configuration features. The fundamental performance features (4 dimensions) focus on the core performance of the system, covering the two original indicators of total energy consumption and total latency. A logarithmic transformation is applied to these indicators to characterize the logarithmic distribution of energy consumption and latency, providing fundamental quantitative information on energy and time dimensions for performance modeling. The task allocation statistical features (6 dimensions) extract multidimensional features for task allocation using statistical methods, including mean, standard deviation, maximum value, and minimum value. The ratio of the maximum to minimum value (after logarithmic transformation) and the coefficient of variation characterize the extreme value ratio and dispersion of task allocation, comprehensively describing the statistical distribution of task allocation. Bandwidth resource allocation statistical characteristics (6 dimensions): Based on bandwidth type resource allocation, normalized mean, standard deviation, maximum, and minimum values are extracted. The extreme value ratio and coefficient of variation are then used to quantify the bandwidth resource allocation characteristics in terms of central tendency, dispersion, and extreme value relationships through logarithmic transformation. System configuration characteristics (4 dimensions): This set of features describes the system's configuration attributes, including the logarithmic transformation of the number of devices and computing power, distance, and heterogeneity of task allocation. SVR (Support Vector Regression) takes a 20-dimensional system state feature vector as input and outputs a predicted reliability score, completing a mapping model from system state to service reliability score.
[0040] Therefore, this invention selects 20-dimensional numerical feature vectors for time-series performance data and uses support vector regression (SVR) to construct a reliability prediction model; by using kernel functions to handle complex nonlinear correlations, it effectively solves the problem of multidimensional evaluation indicators with nonlinear mapping, and significantly improves prediction accuracy.
[0041] In step S102, the initial labeled dataset is augmented based on a preset large language model to obtain an augmented dataset, and the pre-built service reliability score mapping model is trained using the augmented dataset to obtain the trained service reliability score mapping model.
[0042] Specifically, by augmenting the initial labeled dataset, the data bottleneck was overcome, and a service reliability score mapping model that can accurately quantify complex nonlinear relationships was trained. It can accurately learn the complex mapping function from system state features to service reliability scores from the augmented data, solving the challenge of multidimensional evaluation indicators with nonlinear mapping. Finally, a complete solution was built to objectify and quantify subjective user experience, completely changing the traditional dilemma of QoE being difficult to measure and apply, and significantly improving prediction accuracy.
[0043] Optionally, in some embodiments, the initial labeled dataset is augmented based on a preset large language model to obtain an augmented dataset, and a pre-constructed service reliability score mapping model is trained using the augmented dataset to obtain a trained service reliability score mapping model. This includes: determining a validation set based on the initial labeled dataset and constructing a probabilistic reference boundary based on the validation set; injecting Gaussian noise into the multidimensional state feature vector to obtain unlabeled samples; summarizing the nonlinear scoring rules of the original samples in the initial labeled dataset based on the preset large language model to generate reliability scores for the unlabeled samples; performing a quality assessment on the reliability scores according to the probabilistic reference boundary, a preset reasonable interval, and the validation set; if the quality assessment result does not meet preset conditions, iterating the preset large language model based on the quality assessment result to obtain an augmented dataset; generating reliability score labels based on the augmented dataset, and training the service reliability score mapping model based on the reliability score labels.
[0044] Optionally, in some embodiments, the reliability score is evaluated for quality based on a preset reasonable range, including: calculating the precision of the reliability score based on the augmented dataset, calculating the recall of the reliability score based on the validation set; calculating the harmonic mean of the reliability score based on the recall and precision; and determining whether to iterate the harmonic mean of the reliability score based on a preset harmonic mean threshold.
[0045] Among them, recall rate is a measure of the defined reasonable range. Recall is the coverage capability of a data range, meaning what percentage of the actual labeled data points fall within that range. A higher recall indicates a more reasonable range definition, minimizing the omission of truly valid data. Precision (Acc) measures the accuracy of the data selected by the range (i.e., data considered "good"). The harmonic mean (F1-Score) is the harmonic average of recall and precision. The preset harmonic mean threshold can be a user-defined threshold, a threshold obtained through a limited number of experiments, or a threshold obtained through a limited number of computer simulations; no specific limitations are imposed here.
[0046] Understandably, the first step is to partition off a portion of the original dataset (the initial labeled dataset) as a validation set. This validation set is not used for any training; it is only used to define reasonable intervals. Then, a probabilistic reference boundary is constructed based on the initial labeled dataset. This boundary is defined as the range of the mean plus or minus K times the standard deviation. ),in, An adjustable confidence parameter was used to establish an objective criterion for data rationality. Gaussian noise of varying intensities was injected progressively into the seed sample feature vector (multidimensional state feature vector) to generate 410 physically neighboring unlabeled samples. GPT-5 was used for data augmentation. Through a structured contextual learning process, the model first summarizes the original samples during the learning phase. The non-linear scoring pattern is then used to generate reliability scores for unlabeled samples during the labeling stage. The pre-defined reasonable range is considered as the decision boundary of a "binary classifier," which will determine the score of each generated score. The entire validation set is categorized as either "reasonable" or "unreasonable". As "positive examples," the effectiveness of this screening mechanism is evaluated. Based on the quality assessment results, the pre-set large language model is iterated to obtain an augmented dataset. Based on this augmented dataset, reliability score annotations are generated. Defined as the system's service reliability score, used to quantify the quality of user experience (QoE), with a value range of [0, 1]; 1 represents highly reliable, and 0 represents completely unreliable. This score is generated by the function... From the system's underlying feature vectors We learn from (multidimensional state feature vectors): And a service reliability score mapping model is trained based on reliability score annotation.
[0047] Furthermore, the formula for calculating recall is: ; The numerator is the number of samples in the real validation set whose label values fall within the interval [L, U]; the denominator is the total number of samples in the real validation set.
[0048] This invention employs an approximate estimation method based on Bach distance to calculate distribution similarity, as shown in the following formula:
[0049] in, and Let represent the mean of the validation set and the mean of the generated dataset, respectively. and Let each represent its variance.
[0050] The F1 score (harmonic mean) is a comprehensive indicator used to avoid situations where one metric is too high while another is too low, as shown in the following formula: ; Set a harmonic mean threshold α (a preset harmonic mean threshold). If the current harmonic mean is higher than the threshold, the generated data is considered to meet the overall reliability standard and can be used for downstream tasks. If it is lower than the threshold, an optimization mechanism is triggered to implement targeted iteration. Specifically, a low recall indicates that the reasonable range is too strict and the range should be expanded (e.g., by increasing the K value). A low Acc indicates that the quality of the generated data or the range setting is problematic and the LLM prompt process should be optimized or the range parameters should be adjusted. After optimization, the data generation and filtering process is re-executed until the harmonic mean meets the preset harmonic mean threshold.
[0051] Thus, the augmented dataset is obtained. ,in The LLM scoring results were finally combined with the manual and LLM labels to form a training set of 500 results. Used for subsequent SVR training, i.e. This application leverages the powerful contextual semantic understanding capabilities of large language models (LLMs) to automatically generate reliability score labels for unlabeled samples. This method significantly reduces reliance on manual annotation, substantially lowers data acquisition costs, and effectively solves the core bottleneck problem of difficult training data acquisition.
[0052] In step S103, the multidimensional state feature vector is input into the trained service reliability score mapping model, and the trained service reliability score mapping model outputs the service quality assessment service reliability score.
[0053] Specifically, to achieve accurate quantification of service reliability, this application constructs a scoring model based on Support Vector Regression (SVR) (a trained service reliability score mapping model) that maps multidimensional system state indicators to standardized reliability scores in the [0,1] interval. Leveraging SVR's excellent generalization ability on medium-sized datasets and the advantage of kernel functions in handling nonlinear relationships, a mapping model from system state to reliability scores is constructed. Through systematic feature selection, 13 key features are ultimately determined to form the input vector, effectively suppressing overfitting while ensuring validation performance. This model, with its structured feature input, effectively overcomes the dual challenges of lacking evaluation criteria and modeling capabilities, providing interpretable and comparable quantitative evidence for service reliability.
[0054] Therefore, the framework of this application uses service reliability as a measurable agent, proposes a semi-automatic label expansion mechanism based on Large Language Model (LLM), and constructs a mapping relationship from system state to user experience quality (QoE) score, thereby achieving quantitative evaluation. This application integrates system state feature quantification, mapping relationship modeling, and LLM-based training data augmentation techniques, aiming to solve the bottleneck problem of QoE quantification.
[0055] To facilitate a deeper understanding by those skilled in the art of evaluating the large model-driven quality experience of agent-oriented communication networks according to the embodiments of this application, the following is combined with... Figure 2 The embodiments shown will be described in detail.
[0056] Specifically, such as Figure 2 As shown, Figure 2 This is a flowchart of an LLM data augmentation process for a large model-driven quality experience evaluation method for intelligent agent communication networks, according to one embodiment of this application. The manually labeled initial dataset is divided into a training set and a validation set. Based on the reliability scores (expert scores) in the validation set, their mean (μ) and standard deviation (σ) are calculated to construct a probabilistic reference boundary (i.e., the validation interval). The training set is used to teach the LLM through structured prompts, allowing it to summarize the nonlinear scoring rules between system states and reliability scores. Gaussian noise of different intensities is injected into the state feature vectors in the training set to generate a large number of unlabeled samples. The learned LLM is used to generate predicted reliability scores for these unlabeled samples. Batch quality evaluation is performed on the LLM-generated data, calculating three core metrics; the calculated F1 score is compared with a preset threshold (α). Based on the specific metric performance, a targeted optimization mechanism is triggered; the LLM-generated data that passes the quality evaluation is merged with the initial training set to finally form a high-quality augmented dataset.
[0057] Thus, through a rigorous and quantitative quality assessment mechanism, it is ensured that the large language model can controllably produce high-quality and highly reliable training data during the data augmentation process, thereby effectively solving the dual bottlenecks of training data scarcity and difficulty in guaranteeing data quality in QoE evaluation.
[0058] According to the large model-driven quality experience evaluation method for intelligent agent communication networks proposed in this application, this application obtains multi-dimensional state feature vectors and constructs an initial labeled dataset based on the multi-dimensional state feature vectors; it then performs data augmentation on the initial labeled dataset based on a preset large language model to obtain an augmented dataset, and uses the augmented dataset to train a pre-constructed service reliability score mapping model to obtain a trained service reliability score mapping model; the multi-dimensional state feature vectors are input into the trained service reliability score mapping model, and the trained service reliability score mapping model outputs a service quality evaluation service reliability score. Thus, by using a semi-automatic label expansion mechanism based on a large language model, a mapping relationship from system state to user experience quality score is constructed, solving the problem of difficulty in quantifying QoE evaluation and improving evaluation efficiency.
[0059] Next, referring to the accompanying drawings, a large model-driven quality experience evaluation device for intelligent agent communication networks, according to an embodiment of this application, is described.
[0060] Figure 3 This is a block diagram of a large model-driven quality experience evaluation device for intelligent agent communication networks according to an embodiment of this application.
[0061] like Figure 3 As shown, the large model-driven quality experience evaluation device 10 for intelligent agent communication networks includes: an acquisition module 100, a training module 200, and an evaluation module 300.
[0062] The acquisition module 100 is used to acquire multi-dimensional state feature vectors and construct an initial labeled dataset based on the multi-dimensional state feature vectors. The training module 200 is used to perform data augmentation on the initial labeled dataset based on a preset large language model to obtain an augmented dataset, and to use the augmented dataset to train a pre-built service reliability score mapping model to obtain a trained service reliability score mapping model. The evaluation module 300 is used to input multi-dimensional state feature vectors into the trained service reliability score mapping model, and output the service quality evaluation service reliability score through the trained service reliability score mapping model.
[0063] Optionally, the multidimensional state feature vector and the initial labeled dataset include: collecting state samples from multiple intelligent agent devices, wherein the state samples include the latency of the intelligent agent devices, the energy consumption of the intelligent agent devices, the bandwidth allocated to the intelligent agent devices, the computational load of the intelligent agent devices, and the total number of intelligent agent devices; and processing the state samples of multiple intelligent agent devices based on a preset function to obtain a multidimensional state feature vector.
[0064] Optionally, the acquisition module 100 is specifically used for: determining multiple benchmark scenarios; generating a first reliability level state for each benchmark scenario based on a preset efficient optimization strategy, generating a second reliability level state for each benchmark scenario based on a preset suboptimal strategy with perturbation, and generating a third reliability level state for each benchmark scenario based on a preset completely random strategy; acquiring the scoring results generated based on the first reliability level state, the second reliability level state, and the third reliability level state, and obtaining an initial labeled dataset based on the scoring results.
[0065] Optionally, the training module 200 is specifically used for: determining a validation set based on the initial labeled dataset, and constructing a probabilistic reference boundary based on the validation set; injecting Gaussian noise into the multidimensional state feature vector to obtain unlabeled samples; summarizing the nonlinear scoring rules of the original samples in the initial labeled dataset based on a pre-set large language model to generate reliability scores for the unlabeled samples; performing quality assessment on the reliability scores according to the probabilistic reference boundary, a pre-set reasonable interval, and the validation set; if the quality assessment result does not meet the pre-set conditions, iterating the pre-set large language model based on the quality assessment result to obtain an augmented dataset; generating reliability score labels based on the augmented dataset, and training a service reliability score mapping model based on the reliability score labels.
[0066] Optionally, the training module 200 is specifically used for: calculating the precision of the reliability score based on the augmented dataset, calculating the recall of the reliability score based on the validation set; calculating the harmonic mean of the reliability score based on the recall and precision; and determining whether to iterate the harmonic mean of the reliability score based on a preset harmonic mean threshold.
[0067] Optionally, the multidimensional state feature vector includes basic performance features, task allocation statistics, bandwidth resource allocation statistics, and system configuration features.
[0068] It should be noted that the foregoing explanation of the embodiment of the large model-driven quality experience evaluation method for agent communication networks also applies to the large model-driven quality experience evaluation device for agent communication networks in this embodiment, and will not be repeated here.
[0069] According to the large model-driven quality experience evaluation device for intelligent agent communication networks proposed in this application, this application obtains multi-dimensional state feature vectors and constructs an initial labeled dataset based on the multi-dimensional state feature vectors; it then performs data augmentation on the initial labeled dataset based on a preset large language model to obtain an augmented dataset, and uses the augmented dataset to train a pre-constructed service reliability score mapping model to obtain a trained service reliability score mapping model; the multi-dimensional state feature vectors are input into the trained service reliability score mapping model, and the trained service reliability score mapping model outputs a service quality evaluation service reliability score. Thus, by using a semi-automatic label expansion mechanism based on a large language model, a mapping relationship from system state to user experience quality score is constructed, solving the problem of difficulty in quantifying QoE evaluation and improving evaluation efficiency.
[0070] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 401, the processor 402, and the computer program stored on the memory 401 and capable of running on the processor 402.
[0071] When the processor 402 executes the program, it implements the large model-driven quality experience evaluation method for intelligent agent communication networks provided in the above embodiments.
[0072] Furthermore, electronic devices also include: Communication interface 403 is used for communication between memory 401 and processor 402.
[0073] The memory 401 is used to store computer programs that can run on the processor 402.
[0074] Memory 401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0075] If the memory 401, processor 402, and communication interface 403 are implemented independently, then the communication interface 403, memory 401, and processor 402 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 4The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0076] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0077] Processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0078] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described large model-driven quality experience evaluation method for intelligent agent communication networks.
[0079] This application also provides a computer program product that stores a computer program that, when executed by a processor, implements the above-described large model-driven quality experience evaluation method for intelligent agent communication networks.
[0080] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0081] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0082] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0083] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0084] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
Claims
1. A large-model-driven quality experience evaluation method for intelligent agent communication networks, characterized in that, Includes the following steps: Obtain multidimensional state feature vectors, and construct an initial labeled dataset based on the multidimensional state feature vectors; The initial labeled dataset is augmented based on a pre-defined large language model to obtain an augmented dataset, and a pre-built service reliability score mapping model is trained using the augmented dataset to obtain a trained service reliability score mapping model. The multidimensional state feature vector is input into the trained service reliability score mapping model, and the service reliability score is output through the trained service reliability score mapping model.
2. The method according to claim 1, characterized in that, The multidimensional state feature vector and the initial labeled dataset include: Collect state samples of multiple intelligent agent devices, wherein the state samples include the latency of intelligent agent devices, the energy consumption of intelligent agent devices, the bandwidth allocated to intelligent agent devices, the task computation of intelligent agent devices, and the total number of intelligent agent devices; Based on a preset function, the state samples of the multiple intelligent agent devices are processed to obtain the multidimensional state feature vector.
3. The method according to claim 1 or 2, characterized in that, The construction of the initial labeled dataset based on the multidimensional state feature vector includes: Multiple baseline scenarios were identified; Based on a preset efficient optimization strategy, the first reliability level state of each benchmark scenario is generated, and based on a preset suboptimal strategy with perturbation, the second reliability level state of each benchmark scenario is generated, and based on a preset completely random strategy, the third reliability level state of each benchmark scenario is generated. Obtain the scoring results generated based on the first reliability level state, the second reliability level state, and the third reliability level state, and obtain the initial labeled dataset based on the scoring results.
4. The method according to claim 1, characterized in that, The process of augmenting the initial labeled dataset based on a pre-defined large language model to obtain an augmented dataset, and then using the augmented dataset to train a pre-built service reliability score mapping model to obtain a trained service reliability score mapping model, includes: A validation set is determined based on the initial labeled dataset, and a probabilistic reference boundary is constructed based on the validation set; Gaussian noise is injected into the multidimensional state feature vector to obtain unlabeled samples; Based on the preset large language model, the nonlinear scoring rules of the original samples in the initial labeled dataset are summarized to generate a reliability score for the unlabeled samples; The reliability score is evaluated based on the probabilistic reference boundary, the preset reasonable interval, and the validation set. If the quality evaluation result does not meet the preset conditions, the preset large language model is iterated based on the quality evaluation result to obtain the augmented dataset. Based on the augmented dataset, reliability score labels are generated, and the service reliability score mapping model is trained based on the reliability score labels.
5. The method according to claim 4, characterized in that, The quality assessment of the reliability score based on a preset reasonable range includes: The precision of the reliability score is calculated based on the augmented dataset, and the recall of the reliability score is calculated based on the validation set. Calculate the harmonic mean of the reliability score based on the recall rate and the precision rate; Whether to iterate the harmonic mean of the reliability score is determined based on a preset harmonic mean threshold.
6. The method according to claim 1, characterized in that, The multidimensional state feature vector includes basic performance features, task allocation statistics, bandwidth resource allocation statistics, and system configuration features.
7. A large-model-driven quality experience evaluation device for intelligent agent communication networks, characterized in that, include: The acquisition module is used to acquire multidimensional state feature vectors and construct an initial labeled dataset based on the multidimensional state feature vectors. The training module is used to perform data augmentation on the initial labeled dataset based on a preset large language model to obtain an augmented dataset, and to use the augmented dataset to train a pre-built service reliability score mapping model to obtain a trained service reliability score mapping model. The evaluation module is used to input multi-dimensional state feature vectors into the trained service reliability score mapping model, and output a service quality assessment service reliability score through the trained service reliability score mapping model.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the large model-driven quality experience evaluation method for agent communication networks as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the large model-driven quality experience evaluation method for agent communication networks as described in any one of claims 1-6.
10. A computer program product, said computer program product storing a computer program, characterized in that, When executed by the processor, the program implements the large model-driven quality experience evaluation method for agent-oriented communication networks as described in any one of claims 1-6.