User perception optimization method, apparatus, device, medium, and product

By constructing a Markov decision process in an edge computing network to maximize user perception quality, and employing a deep reinforcement learning framework, the dynamic and adaptive issues of user perception optimization methods are addressed, achieving optimization and stability improvement in user perception.

CN122173722APending Publication Date: 2026-06-09CHINA MOBILE GROUP DESIGN INST +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE GROUP DESIGN INST
Filing Date
2024-12-09
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

In existing technologies, user-perceived optimization methods rely on the experience and judgment of professionals, which are time-consuming, costly, slow to respond, and highly subjective. They are difficult to adapt to rapidly changing large-scale network environments. Furthermore, methods based on optimization algorithms mainly focus on static network resource allocation and cannot meet the needs of real-time computing.

Method used

By constructing a user-perceived quality maximization problem based on Markov decision processes, and employing a proximity policy optimization method within a deep reinforcement learning framework, dynamic optimization of user perception is achieved.

Benefits of technology

It achieves accurate description and optimization of user perception in edge computing networks, ensuring the long-term maximization of user experience, adapting to environmental changes, reducing instability during online training, and improving user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173722A_ABST
    Figure CN122173722A_ABST
Patent Text Reader

Abstract

This disclosure relates to a user perception optimization method, apparatus, device, medium, and product in the field of edge computing technology. The method includes: constructing a user perception optimization system model with the goal of enhancing the perception of all users through task processing mode selection; transforming the user perception optimization system model into a user perception quality maximization problem based on Markov decision processes; and employing a proximity policy optimization method within a deep reinforcement learning framework to obtain the optimal task scheduling strategy for user perception in the Markov decision process user perception quality maximization problem. This disclosure establishes a dynamic task scheduling model for user perception in edge computing networks, achieving an accurate description of the user perception problem. Based on the Markov decision process user perception quality maximization problem, it uses a proximity policy optimization method within a deep reinforcement learning framework to intelligently derive the optimal task scheduling strategy, ensuring the perception experience of each user throughout a continuous period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of edge computing technology, and in particular to a user perception optimization method, apparatus, device, medium and product. Background Technology

[0002] With the development of mobile network technology and Internet of Things (IoT) technology, a large number of devices are connected to mobile networks for communication, interaction, and data processing. The enormous demand for data processing has pushed traditional cloud computing to its limits, making it unable to meet users' needs for real-time computing. Therefore, in compute-intensive scenarios, edge computing and caching are used as supplementary methods to overcome the high latency and bandwidth limitations associated with cloud computing, thereby improving user experience. Caching can utilize past computation results to reduce latency and alleviate the burden on computing resources to some extent, but the accuracy of cached result matching is not as good as edge computing. The trade-off between computational tasks and result matching determines user experience and has a significant impact on user satisfaction. Therefore, it is necessary to optimize task execution decisions to improve the quality of user experience and thus increase user satisfaction.

[0003] Current technologies employ traditional manual user-aware optimization methods. These methods rely on the experience and judgment of professionals to optimize networks and services, resulting in drawbacks such as time consumption, high cost, slow response, strong subjectivity, difficulty in scaling, insufficient data utilization, and susceptibility to errors. They are also ill-suited to rapidly changing large-scale network environments. Other traditional manual user-aware optimization methods increase system complexity and incur high upfront investment and operating costs. Finally, there are user-aware optimization methods based on optimization algorithms, which primarily focus on the static allocation of network resources. Summary of the Invention

[0004] To address the aforementioned technical problems, embodiments of this disclosure provide a user perception optimization method, apparatus, device, medium, and product to achieve dynamic optimization of user perception.

[0005] A first aspect of this disclosure provides a user-perception optimization method, comprising:

[0006] To enhance the perception of all users by selecting task processing methods, a user perception optimization system model is constructed. The task processing methods include: calculating tasks and matching cached results.

[0007] The user perception optimization system model is transformed into a user perception quality maximization problem based on Markov decision processes.

[0008] In the problem of maximizing user-perceived quality in the Markov decision process, a neighbor policy optimization method based on a deep reinforcement learning framework is used to obtain the user-perceived optimal task scheduling strategy.

[0009] A second aspect of this disclosure provides an apparatus comprising:

[0010] The building module is configured to construct a user perception optimization system model with the goal of enhancing the perception of all users through the selection of task processing methods, wherein the task processing methods include: computing tasks and matching cached results;

[0011] The transformation module is configured to transform the user perception optimization system model into a user perception quality maximization problem based on Markov decision processes.

[0012] The optimization module is configured to employ a deep reinforcement learning framework-based proximity policy optimization method to obtain the user-perceived optimal task scheduling strategy in the user-perceived quality maximization problem of the Markov decision process.

[0013] A third aspect of this disclosure provides an electronic device, including:

[0014] At least one processor;

[0015] Memory for storing the at least one processor-executable instruction;

[0016] The at least one processor is used to execute the instructions to implement the above-described method.

[0017] A fourth aspect of this disclosure provides a non-transient computer-readable storage medium that, when instructions in the non-transient computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the methods described above.

[0018] A fifth aspect of this disclosure provides a computer program product, including a computer program that, when executed, implements the method described above.

[0019] The at least one technical solution adopted in this disclosure can achieve the following beneficial effects: A dynamic task scheduling model for user perception in edge computing networks is established, taking into account the randomness, continuity, and dynamism of user tasks, thus accurately describing the user perception problem. To ensure the long-term maximization of user perception experience, the user perception optimization system model is transformed into a user perception quality maximization problem based on Markov decision processes. A deep reinforcement learning framework, proximity policy optimization, is used to intelligently derive the optimal task scheduling strategy, considering not only the user perception experience at the current moment but also the user perception experience at future moments, ensuring the perception experience of each user throughout the continuous period. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0021] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a user-perception optimization method provided in an embodiment of this disclosure;

[0023] Figure 2 A schematic diagram illustrating the implementation steps of a user-perception optimization method provided in this embodiment of the disclosure;

[0024] Figure 3 A schematic diagram of the structure of a user-perception optimization device provided in an embodiment of this disclosure;

[0025] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;

[0026] Figure 5 This is a schematic diagram of the structure of an exemplary computer system provided in an embodiment of the present disclosure. Detailed Implementation

[0027] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0028] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0029] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The following is combined Figures 1-5 This disclosure describes the user perception optimization methods, apparatus, devices, media, and products provided in the embodiments of this disclosure.

[0033] Figure 1 This is a flowchart illustrating a user-perception optimization method provided in an embodiment of this disclosure, as shown below. Figure 1 As shown in the embodiments of this disclosure, a user-perception optimization method includes:

[0034] S101. To enhance the perception of all users by selecting task processing methods, a user perception optimization system model is constructed. The task processing methods include: calculating tasks and matching cached results.

[0035] Because user tasks in edge computing scenarios are characterized by continuity and dynamism, by constructing a dynamic task scheduling model for user perception in 5G edge computing networks, it is possible to respond to environmental changes and system status in real time, and continuously improve user perception through continuous optimization.

[0036] S102. Transform the user perception optimization system model into a user perception quality maximization problem based on Markov decision process.

[0037] Compared to cloud computing, edge computing servers have limited computing resources. By taking into account the impact of past moments, the original dynamic optimization problem is transformed into a Markov decision process, which can predict the impact of current decisions on the future. Considering the entire time series helps to make globally optimal decisions.

[0038] S103. In the problem of maximizing user-perceived quality in the Markov decision process, a neighbor policy optimization method based on a deep reinforcement learning framework is used to obtain the user-perceived optimal task scheduling strategy.

[0039] To address the challenges of high environmental randomness caused by the diversity of tasks in edge computing scenarios, this paper proposes a deep reinforcement learning framework based on the neighbor strategy. By considering sample efficiency, it can quickly learn effective strategies with limited data. Furthermore, by limiting the magnitude of strategy updates, it reduces instability during online training, updates model parameters in real time, ensures the acquisition of the optimal task scheduling strategy, and safeguards user experience.

[0040] In the embodiments of this disclosure, due to the continuous and dynamic nature of user tasks in edge computing scenarios, a dynamic task scheduling model for user perception in 5G edge computing networks is constructed. This model can respond to environmental changes and system status in real time and continuously improve user perception through optimization. Compared to cloud computing, edge computing servers have limited computing resources. Considering the impact of past moments, the original dynamic optimization problem is transformed into a Markov decision process, which can predict the impact of current decisions on the future. Considering the entire time series helps to make globally optimal decisions. In the face of the strong environmental randomness caused by the diversity of tasks in edge computing scenarios, based on the proximity policy of the deep reinforcement learning framework, considering sample efficiency, effective policies can be learned quickly with limited data. By limiting the magnitude of policy updates, the instability during online training is reduced, and the model parameters are updated in real time to ensure that the optimal task scheduling policy is obtained, thus guaranteeing user perception.

[0041] In one embodiment of this disclosure, Figure 2 This is a schematic diagram illustrating the implementation steps of a user-perception optimization method provided in an embodiment of this disclosure, as shown below. Figure 2 As shown, this disclosure constructs a user-aware dynamic task scheduling model in a 5G edge computing network. Under computational resource constraints, the dynamic task scheduling model is transformed into a user-aware quality maximization problem based on a Markov decision process. Considering the uncertainty and dynamism of the state space, a user-aware optimization algorithm based on deep reinforcement learning is introduced, employing an autonomous intelligent decision-making scheme between task computation and cache matching to handle continuously incoming user tasks.

[0042] S1, User Perception Optimization System Model.

[0043] For example, consider a 5G network with edge computing capabilities. Edge servers are deployed at the central base station, and these servers cache a large amount of historical data to support task computation and downloading cached results. In this network, a large number of users have computationally intensive tasks to process. Since local computing resources are insufficient to support these tasks, they need to be offloaded to edge servers for processing. Due to the large number of users and the limited computing resources of the edge servers, the edge servers will choose whether to perform computation on the task or match a similar cached result.

[0044] Considering the practical situation, tasks from different users will enter the edge server sequentially in a time-order manner. For all tasks, a time-slot model is established as: t = {1, 2, ..., t, ...}, where the start time of each time slot is the time the task enters, and the end time is the time the next task enters the edge server. Since the generation times of tasks from different users differ, the duration of each time slot also varies. The task generated by a user in each time slot can be represented as a tuple Q(t) = {T...} s (t),w(t),d ca (t),p ca (t),α(t)}, where T s w(t) represents the time when the task starts entering the server, w(t) represents the computing resources required to complete the task, and d ca (t) represents the time when the task matches the cached result, p ca It represents the accuracy of the cached results, and α(t) represents the user's tolerance for the accuracy of the task.

[0045] When an edge server selects to perform computation on a task, it allocates a fixed amount of computing resources f to each task. Since edge server computing resources are limited, the duration of the computing resources allocated to each task is T. max Resources will be reclaimed after the time limit is exceeded to ensure the smooth execution of subsequent computing tasks.

[0046] For the computation task at time t, if the edge server chooses to allocate computing resources for computation, the time required to complete the computation can be calculated as follows:

[0047]

[0048] Subject to resource recovery time T max Due to limitations, the actual computation latency for each task is corrected to:

[0049]

[0050] If computing resources are reclaimed before a task is completed, the accuracy of the task results will be compromised. The accuracy of the task calculation results can be expressed as:

[0051]

[0052] User perception is closely related to the accuracy of task results and processing latency. Generally speaking, performing tasks using computation methods yields higher accuracy but incurs greater computational latency, while directly matching cached results offers lower latency but accuracy is difficult to guarantee. Therefore, when edge servers execute tasks, a trade-off between accuracy and latency must be struck when choosing between computational tasks and matching cached results. User perception based on accuracy can be calculated as follows:

[0053] p(t)=x(t)p co (t)+(1-x(t))p ca (t)

[0054] Where x(t) is a binary variable, that is,

[0055]

[0056] To balance the impact of latency and accuracy on user perception, it is necessary to normalize the latency. Since the tasks involved in this embodiment are computationally intensive rather than data-intensive, with extremely small data volumes, and considering the extremely high speeds of 5G networks, the transmission and result return latency of the tasks can be ignored. Therefore, the user perception based on task processing latency can be calculated as follows:

[0057]

[0058] Therefore, the user's perception of combining task result accuracy and task processing delay can be expressed as:

[0059] U(t)=α(t)p(t)+(1-α(t))T(t)

[0060] Due to limited computing resources, it is necessary to ensure the availability of remaining computing resources before allocating computing resources to tasks. The remaining computing resources in the current time slot can be calculated as follows:

[0061]

[0062] Where F represents the total computing resources of the edge server, and ρ(j) represents whether the computing resources allocated to the task in time slot j have been reclaimed, it can be determined as follows:

[0063]

[0064] The objective of this disclosed solution is to enhance the user experience for all users by selecting appropriate task processing methods. Specifically, the user experience optimization system model is constructed based on user experience based on task result accuracy and task processing latency. This involves determining a dynamic perception optimization problem over an infinite time frame and defining constraints. These constraints include a first constraint and a second constraint. The first constraint states that the remaining computing resources in the current time slot are greater than or equal to the fixed computing resources allocated to any task. The second constraint states that edge servers can only choose to compute tasks or match cached results. Specifically, for each UE, the decision is made whether to compute a task or match cached results in time slot t. Considering the large number of users involved, the following dynamic optimization problem over an infinite time frame is proposed:

[0065]

[0066] Here, C1 ensures that the computing resources allocated to all decision tasks at any given time do not exceed the total computing resources of the edge server; C2 indicates that the server can only select computing tasks or match cached results.

[0067] S2. User Perception Optimization Method Based on Deep Reinforcement Learning

[0068] Regarding problem P1, the choice of task processing method can be considered a dynamic sequence problem. The state at time slot t+1 depends only on the state and action at time slot t, and is independent of the task selection decisions before time slot t. Therefore, it exhibits the Markov property. In the technical solution disclosed in this paper, problem P1 is transformed into a Markov decision (MDP) problem with an infinite time range, and a deep reinforcement learning-based algorithm is proposed to solve it.

[0069] S2.1 User Perception Optimization Problem Based on MDP

[0070] In this embodiment, the transformed MDP problem consists of a triple M = {S, A, R}, where S represents the state space, A represents the action space, and R represents the reward function. In each time slot t, the edge server can act as an agent to perceive the current state s(t) ∈ S of the 5G edge computing network system and select an action a(t) ∈ A to execute. After the action is completed, the system state is updated to s(t+1), and the agent receives a reward from the environmental feedback.

[0071] The MDP model designed in detail in this embodiment is shown below:

[0072] (1) Define the state space S: When making decisions about the tasks to be performed in time slot t, it is necessary to collect all environmental information of the current state of the 5G edge computing network system and define it as the state space s(t) of time slot t, expressed as:

[0073] s(t)={T s (t),w(t),d ca (t),p ca (t),α(t),f,T max ,F(t)}

[0074] (2) Define the action space A: Based on the current status s(t) of the 5G edge computing network system, the edge server will make the following decisions:

[0075] a(t) = {0, 1}

[0076] The decision is subject to constraint C1. That is, the decision made by the edge server is subject to the first constraint condition, which is: the remaining computing resources in the current time slot are greater than or equal to the fixed computing resources allocated to any task.

[0077] (3) Define the reward function R: Generally, the goal of the MDP problem is to optimize the long-term expectation of the system. For problem P1, this is achieved by optimizing the user perception for each time slot t as the immediate reward function of the edge computing network system, ensuring that the expected user perception is maximized across all time slots, i.e.

[0078]

[0079] The reward function proposed in this embodiment can effectively avoid the sparsity of rewards, enabling edge servers to explore the optimal strategy more quickly.

[0080] S2.2 Design of User Perception Optimization Algorithm Based on Deep Reinforcement Learning

[0081] In the infinite MDP problem with a constantly changing environment, derived from problem P1, this scheme employs the deep reinforcement learning framework Proximity Policy Optimization (PPO) to derive the optimal task scheduling policy. Based on the policy gradient method, it aims to train the policy within the reinforcement learning setting to maximize the expected cumulative reward.

[0082] That is, step S103 includes:

[0083] S1031. Initialize the policy network and value function network;

[0084] Before training begins, two parts need to be initialized: the policy network and the value function network. The policy is a function that outputs actions based on the state, while the value function is a function that outputs values ​​based on the state.

[0085] S1032. Obtain the environmental state and select an action according to the strategy;

[0086] In the PPO framework, the policy network represents the probability a(t) of taking an action given a state s(t) for the task decision output. The policy network is typically represented as:

[0087] π(a(t)|s(t);θ)

[0088] Where θ represents the parameters of the policy network.

[0089] The PPO framework combines a value function network, which estimates the state-value function, with a value function network for advantage calculation. This value function network is also a deep neural network, and can be represented as follows: in The parameters represent the values ​​of a value function network. In deep learning, value function networks are allowed to estimate state values ​​to better compute advantages.

[0090] S1033, Perform the action and obtain the environmental reward and the next state;

[0091] S1034. Calculate the advantage estimation function and action value function;

[0092] To measure the improvement in reward for each task processing decision relative to the average level, it is necessary to estimate the advantage of the action. The advantage estimation function is calculated as follows:

[0093] a(s(t),a(t))=Q(s(t),a(t))-V(s(t))

[0094] Where Q(s(t), a(t)) represents the action-value function of the state-action pair, and V(s(t)) represents the state-value function, which can be estimated using a value function network. According to the Bellman equation, Q(s(t), a(t)) can be expressed as:

[0095]

[0096] S1035. Calculate the value loss function and the proxy loss function;

[0097] In problem P1, which includes the Bellman equation, γ represents the discount factor, indicating the impact of future rewards on the decision action executed in the current time slot t. The value loss function is calculated using mean squared error as a metric during the training of the value function network.

[0098]

[0099] The policy is optimized by defining a proxy loss function, which can be expressed as:

[0100]

[0101] Where π represents the probability distribution of the new strategy, π old Let θ be the probability distribution of the old policy, A be the advantage of the action, and ε be a hyperparameter used to control the range of the policy agent loss clipping. Gradient ascent is used to update the parameters θ of the policy network and the parameters of the value function network by maximizing the agent loss function. This allows for improvements to the strategy and value function network to maximize expected returns.

[0102] S1036, Backpropagation, updating the policy network and value function network;

[0103] Repeat steps S1032-S1036 until the stopping condition is met.

[0104] In the embodiments of this disclosure, due to the continuous and dynamic nature of user tasks in edge computing scenarios, the technical solution of this disclosure constructs a dynamic task scheduling model for user perception in 5G edge computing networks. This model can respond to environmental changes and system status in real time and continuously improve user perception through continuous optimization. Compared to cloud computing, edge computing servers have limited computing resources. The technical solution of this disclosure considers the influence of past moments and transforms the original dynamic optimization problem into a Markov decision process, which can predict the impact of current decisions on the future. Considering the entire time series helps to make globally optimal decisions. Facing the problem of strong environmental randomness caused by the diversity of tasks in edge computing scenarios, the technical solution of this disclosure is based on the proximity strategy of the deep reinforcement learning framework. Considering sample efficiency, it can quickly learn effective strategies with limited data. By limiting the magnitude of strategy updates, it reduces the instability during online training and updates model parameters in real time to ensure that the optimal task scheduling strategy is obtained, thus guaranteeing user perception.

[0105] The above description is only a preferred embodiment of this disclosure. It should be noted that, for those skilled in the art, several improvements, optimizations and modifications can be made without departing from the principles of this disclosure, and these should also be considered within the protection scope of this disclosure.

[0106] Figure 3 This is a schematic diagram of the structure of a user-perception optimization device provided in an embodiment of the present disclosure, as shown below. Figure 3 As shown, the device 300 includes:

[0107] Module 301 is configured to build a user perception optimization system model with the goal of enhancing the perception of all users through the selection of task processing methods, wherein the task processing methods include: computing tasks and matching cached results;

[0108] The transformation module 302 is configured to transform the user perception optimization system model into a user perception quality maximization problem based on Markov decision processes.

[0109] The optimization module 303 is configured to use a deep reinforcement learning framework neighbor policy optimization method to obtain the user-perceived optimal task scheduling strategy in the user-perceived quality maximization problem of the Markov decision process.

[0110] In some embodiments, the construction module 301 is configured to determine a dynamic perception optimization problem over an infinite time range based on user perception of task result accuracy and user perception of task processing latency, and to determine constraints.

[0111] In some embodiments, the constraints include: a first constraint and a second constraint, wherein the first constraint is that the remaining computing resources in the current time slot are greater than or equal to the fixed computing resources allocated to any task; and the second constraint is that the edge server can only select computing tasks or match cached results.

[0112] In some embodiments, the user's perception of the accuracy of the task results is determined by the accuracy of the task calculation results and the accuracy of the cached results.

[0113] In some embodiments, the user perception of task processing delay is determined by resource reclamation time, the actual computational latency of each task, and the time it takes for a task to match a cached result.

[0114] In some embodiments, the conversion module 302 includes:

[0115] The state space module is configured to define the state space: collect all environmental information about the current state of the edge computing network system and define it as the state space of time slot t;

[0116] The action space module is configured to define the action space: decisions made by the edge server include: computation tasks; matching cached results;

[0117] The reward function module is configured to define a reward function: optimizing the user perception for each time slot as the immediate reward function for the edge computing network system.

[0118] In some embodiments, the decisions made by the edge server are subject to a first constraint, which is that the remaining computing resources in the current time slot are greater than or equal to the fixed computing resources allocated to any task.

[0119] In some embodiments, the optimization module 303 includes:

[0120] The first module is configured to execute S1031, initialize the policy network and the value function network;

[0121] The second module is configured to execute S1032, obtain the environment status, and select an action according to the policy;

[0122] The third module is configured to execute S1033, perform actions, and obtain environmental rewards and the next state;

[0123] The fourth module is configured to execute S1034, calculate the advantage estimation function and the action value function;

[0124] The fifth module is configured to execute S1035, calculate the value loss function and the proxy loss function;

[0125] The sixth module is configured to execute S1036, backpropagation, and update the policy network and value function network;

[0126] The seventh module is configured to repeatedly execute steps S1032-S1036 until the stopping condition is met.

[0127] In some embodiments, the policy network outputs the probability of taking action for a task decision given a state.

[0128] In some embodiments, the advantage estimation function is determined by an action value function and a state value function representing a state-action pair.

[0129] In some embodiments, mean squared error is used as a metric when calculating the value loss function.

[0130] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 4 As shown, this disclosure also provides an electronic device 400, which includes at least one processor 401 and a memory 402 coupled to the processor 401. The memory 402 is used to store at least one processor 401 executable instructions, wherein the at least one processor 401 is used to execute the instructions to implement the steps of the method described above in this disclosure.

[0131] The processor 401 described above can also be referred to as a Central Processing Unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method described in this embodiment can be implemented by the integrated logic circuitry in the hardware of the processor 401 or by instructions in software form. The processor 401 described above can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method in conjunction with this embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 402, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 401 reads information from the memory 402 and, in conjunction with its hardware, completes the steps of the method described above.

[0132] Figure 5This is a schematic diagram of an exemplary computer system provided by an embodiment of the present disclosure. Various operations / processes according to embodiments of the present disclosure, implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, for example... Figure 5 The computer system 500 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above.

[0133] Computer system 500 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of this disclosure described and / or claimed herein.

[0134] like Figure 5 As shown, the computer system 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the computer system 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0135] Multiple components in the computer system 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device capable of inputting information into the computer system 500. The input unit 506 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 508 may include, but is not limited to, a hard disk and an optical disk. The communication unit 509 allows the computer system 500 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, a modem, network card, infrared communication device, wireless communication transceiver, and / or chipset, such as Bluetooth™ devices, Wi-Fi devices, WiMax devices, cellular communication devices, and / or the like.

[0136] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the methods described in the embodiments of this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 502 and / or communication unit 509. In some embodiments, the computing unit 501 can be configured to perform the methods described in the embodiments of this disclosure by any other suitable means (e.g., by means of firmware).

[0137] This disclosure provides a non-transient computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the methods described in this disclosure.

[0138] Non-transient computer-readable storage media can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0139] It should be noted that the non-transient computer-readable storage medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), or any suitable combination thereof.

[0140] Embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the uplink signal equalization method described above.

[0141] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0142] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.

[0143] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.

[0144] It should be noted that, in this document, terms such as "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0145] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A user-perception optimization method, characterized in that, include: To enhance the perception of all users by selecting task processing methods, a user perception optimization system model is constructed. The task processing methods include: calculating tasks and matching cached results. The user perception optimization system model is transformed into a user perception quality maximization problem based on Markov decision processes. In the problem of maximizing user-perceived quality in the Markov decision process, a neighbor policy optimization method based on a deep reinforcement learning framework is used to obtain the user-perceived optimal task scheduling strategy.

2. The method according to claim 1, characterized in that, The goal of enhancing user perception for all users through the selection of task processing methods, and the construction of a user perception optimization system model, includes: Based on user perceptions of task result accuracy and task processing latency, a dynamic perception optimization problem is identified over an infinite time range, and constraints are determined.

3. The method according to claim 2, characterized in that, The constraints include a first constraint and a second constraint. The first constraint is that the remaining computing resources in the current time slot are greater than or equal to the fixed computing resources allocated to any task. The second constraint is that the edge server can only select computing tasks or match cached results.

4. The method according to claim 2, characterized in that, The user's perception of the accuracy of the task results is determined by the accuracy of the task calculation results and the accuracy of the cached results.

5. The method according to claim 2, characterized in that, The user's perception of task processing delay is determined by the resource reclamation time, the actual computational latency of each task, and the time it takes for the task to match the cached result.

6. The method according to any one of claims 1-5, characterized in that, The process of transforming the user-perceived optimization system model into a user-perceived quality maximization problem based on Markov decision processes includes: Define the state space: Collect all environmental information about the current state of the edge computing network system and define it as the state space of time slot t; Defining the action space: The decisions made by the edge server include: computing tasks; matching cached results; Define the reward function: optimize the user perception for each time slot as the immediate reward function for the edge computing network system.

7. The method according to claim 6, characterized in that, The decisions made by the edge server are subject to a first constraint: the remaining computing resources in the current time slot are greater than or equal to the fixed computing resources allocated to any task.

8. The method according to claim 7, characterized in that, In the problem of maximizing user-perceived quality in the Markov decision process, the nearest neighbor policy optimization method using a deep reinforcement learning framework is employed to obtain the user-perceived optimal task scheduling strategy, including: S1031. Initialize the policy network and value function network; S1032. Obtain the environmental state and select an action according to the strategy; S1033, Perform the action and obtain the environmental reward and the next state; S1034. Calculate the advantage estimation function and action value function; S1035. Calculate the value loss function and the proxy loss function; S1036, Backpropagation, updating the policy network and value function network; Repeat steps S1032-S1036 until the stopping condition is met.

9. The method according to claim 8, characterized in that, The policy network outputs the probability of taking action for task decision given a state.

10. The method according to claim 8, characterized in that, The advantage estimation function is determined by the action value function and the state value function representing the state-action pair.

11. The method according to claim 8, characterized in that, When calculating the value loss function, the mean square error is used as a metric.

12. An apparatus, characterized in that, include: The building module is configured to construct a user perception optimization system model with the goal of enhancing the perception of all users through the selection of task processing methods, wherein the task processing methods include: computing tasks and matching cached results; The transformation module is configured to transform the user perception optimization system model into a user perception quality maximization problem based on Markov decision processes. The optimization module is configured to employ a deep reinforcement learning framework-based proximity policy optimization method to obtain the user-perceived optimal task scheduling strategy in the user-perceived quality maximization problem of the Markov decision process.

13. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-11.

14. A non-transient computer-readable storage medium, characterized in that, When the instructions in the non-transient computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-11.

15. A computer program product, comprising a computer program, characterized in that, When executed, the computer program implements the method as described in any one of claims 1-11.