Communication method and device

By deploying a reward model in communication devices and updating the prior probabilities of actions, the problem of insufficient decision-making accuracy and reliability in existing technologies is solved, enabling a more efficient and personalized decision-making process.

CN122073739APending Publication Date: 2026-05-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-11-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing communication equipment relies on predefined action probabilities for decision-making, resulting in poor accuracy and reliability of the decisions.

Method used

By deploying a reward model, reward information for actions is obtained based on the reward model, and the prior probabilities of network devices are updated to improve the accuracy and reliability of decision-making.

Benefits of technology

It improves the effectiveness and reliability of decision-making, adapts to the personalized needs of different users, and reduces computational load and signaling overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073739A_ABST
    Figure CN122073739A_ABST
Patent Text Reader

Abstract

The invention provides a communication method and device, and the method comprises the steps: determining rewards corresponding to K actions based on a reward model, and K is a positive integer; and sending first information to a first network device, the first information being used for indicating rewards corresponding to the K actions, and the first information being used for updating prior probabilities corresponding to the K actions. The first network device obtains which of the actions (or actions) can obtain a higher reward, so that the first network device can make a better decision. Moreover, the award is higher in accuracy compared with the prior probability, so that the decision making based on the updated probability is higher in effectiveness and reliability compared with the decision making based on the prior probability corresponding to each action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communications, and more particularly to a communication method and apparatus in the field of communications. Background Technology

[0002] Optimization of decision-making processes is an important research area, especially in the application of machine learning and artificial intelligence. How to effectively select and evaluate actions to achieve optimal decisions is a key issue.

[0003] Currently, communication devices can rely on predefined probabilities corresponding to each action to guide the decision-making process. However, the accuracy of these probabilities is relatively poor, which affects the reliability of the decisions made. Summary of the Invention

[0004] This application provides a communication method and apparatus to improve the reliability of decision-making.

[0005] Firstly, this application provides a communication method, which can be executed by a first device. The first device can be a terminal device, or a circuit or chip applicable to a terminal device (such as a modem chip, also known as a baseband chip, or a system-on-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip), and this application does not limit its scope. The first device can also be a digital twin (DT), which can be deployed in network equipment or on other devices, and this application does not limit its scope.

[0006] For example, the method includes: determining the rewards corresponding to K actions based on a reward model, where K is a positive integer; sending first information to a first network device, the first information being used to indicate the rewards corresponding to the K actions, and the first information being used to update the prior probabilities corresponding to the K actions.

[0007] In the above scheme, the first device can deploy a reward model. Based on the reward model, the reward corresponding to each action is obtained, and the reward corresponding to the action is indicated to the first network device. This allows the first network device to update the prior probabilities corresponding to the actions. By indicating the reward for each action, the first network device can determine which actions (or actions) can obtain higher rewards, enabling it to make better decisions. Moreover, the rewards corresponding to the actions are more accurate than the predefined prior probabilities. Therefore, making decisions based on the updated probabilities is more effective and reliable than making decisions based on the prior probabilities corresponding to each action.

[0008] In conjunction with the first aspect, in some possible implementations of the first aspect, the method further includes: receiving second information from the first network device, the second information being used to indicate the K actions described above.

[0009] In other words, the first network device can instruct the first device which actions' corresponding rewards need to be obtained. In this way, the first device can obtain the rewards corresponding to these actions based on the reward model and then instruct the first network device.

[0010] Optionally, the above K actions are K of one or more actions generated based on the Monte Carlo Tree Search (MCTS) model.

[0011] In conjunction with the first aspect, in some possible implementations of the first aspect, the method further includes: receiving third information from the first network device, the third information being used to indicate the first path; and sending a score corresponding to the first path to the first network device.

[0012] Optionally, the above score is a weighted average of the environment score of the first path and the reward of the first path obtained based on the reward model.

[0013] In the above scheme, the scores can be more comprehensive and accurate by taking into account both the environmental score and the reward obtained based on the reward model, thereby improving the reliability and accuracy of the decision-making process.

[0014] In conjunction with the first aspect, in some possible implementations of the first aspect, when the above method is applied to the first network device for beam prediction, the above K actions include the K candidate beams predicted by the first network device.

[0015] In conjunction with the first aspect, in some possible implementations of the first aspect, the rewards corresponding to the above K actions are obtained based on the reward model and the measurement results corresponding to the K candidate beams.

[0016] In conjunction with the first aspect, in some possible implementations of the first aspect, the aforementioned first information is carried in channel state information (CSI).

[0017] In conjunction with the first aspect, in some possible implementations of the first aspect, when the above method is applied to the first network device for resource scheduling, the above K actions include K candidate scheduling users predicted by the first network device.

[0018] In conjunction with the first aspect, in some possible implementations of the first aspect, the method further includes: receiving fourth information from the first network device, the fourth information indicating K first probabilities corresponding to the K actions, the K first probabilities being updated prior probabilities corresponding to the K actions associated with the first terminal, or the K first probabilities being obtained based on updated prior probabilities associated with multiple terminals, the multiple terminals including the first terminal, each of the multiple terminals being associated with updated prior probabilities corresponding to the K actions; and indicating the K first probabilities to the second network device.

[0019] In other words, the first network device can send the first probabilities corresponding to the aforementioned K actions to the first device, so that when another network device, such as the second network device, makes a decision, the first device can directly report the first probabilities corresponding to the aforementioned K actions to the second network device. The second network device does not need to request the rewards corresponding to the aforementioned K actions from the first device to update its stored prior probabilities for the aforementioned K actions, and the first device also does not need to calculate the rewards corresponding to the aforementioned K actions. This not only reduces the computational load on the first device but also reduces signaling overhead.

[0020] Secondly, this application provides another communication method, which includes: receiving first information from a first device, the first information being used to indicate rewards corresponding to K actions, the rewards corresponding to the K actions being obtained based on a reward model, where K is a positive integer; and updating the prior probabilities corresponding to the K actions based on the first information.

[0021] In one possible implementation, the method is performed by a first network device. The first network device may be the first network device itself, or it may be a circuit or chip applicable to the first network device, etc., and this application does not limit it in this regard.

[0022] In conjunction with the second aspect, in some possible implementations of the second aspect, the method further includes: sending second information to the first device, the second information being used to instruct the aforementioned K actions.

[0023] Optionally, the above K actions are K of one or more actions generated based on the MCTS model.

[0024] In conjunction with the second aspect, in some possible implementations of the second aspect, the above method further includes: sending third information to the first device, the third information being used to indicate the first path; and receiving a score corresponding to the first path.

[0025] In conjunction with the second aspect, in some possible implementations of the second aspect, the above score is a weighted average of the environment score of the first path and the reward of the first path obtained based on the reward model.

[0026] In conjunction with the second aspect, in some possible implementations of the second aspect, when the above method is applied to the first network device for beam prediction, the above K actions include the K candidate beams predicted by the first network device.

[0027] In conjunction with the second aspect, in some possible implementations of the second aspect, the rewards corresponding to the aforementioned K actions are obtained based on the reward model and the measurement results corresponding to the K candidate beams.

[0028] In conjunction with the second aspect, in some possible implementations of the second aspect, the aforementioned first information is carried in the CSI.

[0029] In conjunction with the second aspect, in some possible implementations of the second aspect, when the above method is applied to the first network device for resource scheduling, the above K actions include the K candidate scheduling users predicted by the first network device.

[0030] In conjunction with the second aspect, in some possible implementations of the second aspect, the method further includes: sending fourth information to the first device, the fourth information indicating K first probabilities corresponding to the K actions, the K first probabilities being updated prior probabilities corresponding to the K actions associated with the first terminal, or the K first probabilities being obtained based on updated prior probabilities associated with multiple terminals, the multiple terminals including the first terminal, each of the multiple terminals being associated with the updated prior probabilities corresponding to the K actions.

[0031] Thirdly, this application provides a communication device for executing the methods in the first aspect and any possible implementation thereof, or for executing the methods in the second aspect and any possible implementation thereof. Specifically, the communication device includes a module for executing the methods in the first aspect and any possible implementation thereof, or includes a module for executing the methods in the second aspect and any possible implementation thereof.

[0032] Fourthly, this application provides another communication device, including a processor coupled to a memory, which can be used to execute instructions in the memory to implement the methods in the first aspect and any possible implementation thereof, or to implement the methods in the second aspect and any possible implementation thereof. Optionally, the communication device further includes a memory. Optionally, the communication device further includes a communication interface, and the processor is coupled to the communication interface.

[0033] In one implementation, the communication device is a terminal device or a network device. When the communication device is a terminal device or a network device, the communication interface can be a transceiver, or an input / output interface.

[0034] In another implementation, the communication device is a chip applicable to terminal devices or network devices. When the communication device is a chip applicable to terminal devices or network devices, the aforementioned communication interface can be an input / output interface.

[0035] Fifthly, a processor is provided, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is configured to receive signals through the input circuit and transmit signals through the output circuit, causing the processor to execute the methods of the first aspect and any possible implementation thereof, or to execute the methods of the second aspect and any possible implementation thereof.

[0036] In the specific implementation process, the processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, gate circuit, flip-flop, and various logic circuits. The input signal received by the input circuit can be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit can be, for example, but not limited to, output to a transmitter and transmitted by the transmitter. Furthermore, the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. This application does not limit the specific implementation method of the processor and various circuits.

[0037] In a sixth aspect, a communication device is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory, receive signals via a receiver, and transmit signals via a transmitter to execute the methods described in the first aspect and any possible implementation thereof, or to execute the methods described in the second aspect and any possible implementation thereof.

[0038] Optionally, the processor may be one or more, and the memory may be one or more.

[0039] Optionally, the memory may be integrated with the processor, or the memory may be separated from the processor.

[0040] In the specific implementation process, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same chip or set on different chips. This application does not limit the type of memory or the way the memory and processor are set.

[0041] It should be understood that related data interaction processes, such as sending information, can be seen as a process of the processor outputting information, and receiving information can be seen as a process of the processor receiving input information. Specifically, the data output by the processor can be sent to the transmitter, and the input data received by the processor can come from the receiver. The transmitter and receiver can be collectively referred to as a transceiver.

[0042] The communication device in the sixth aspect above can be a chip. The processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. The memory can be integrated into the processor or located outside the processor and exist independently.

[0043] In a seventh aspect, a computer program product is provided, the computer program product comprising: a computer program (also referred to as code or instructions), which, when the computer program is run, causes a computer to perform the methods of the first aspect and any possible implementation thereof, or to perform the methods of the second aspect and any possible implementation thereof.

[0044] Eighthly, a computer-readable storage medium is provided that stores a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the methods of the first aspect and any possible implementation thereof, or to perform the methods of the second aspect and any possible implementation thereof.

[0045] It should be understood that the second to eighth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of a communication system applied in an embodiment of this application;

[0047] Figure 2 This is a schematic diagram of existing knowledge distillation;

[0048] Figure 3 This is a schematic diagram of the MCTS process provided in the embodiments of this application;

[0049] Figure 4 This is a comparative schematic diagram of prior art and the method of this application provided in the embodiments of this application;

[0050] Figure 5This is a flowchart illustrating the communication method provided in an embodiment of this application;

[0051] Figure 6 This is a schematic diagram of the interaction between the first terminal device and the second network device provided in the embodiments of this application;

[0052] Figure 7 This is a schematic diagram of the beam prediction process provided in an embodiment of this application;

[0053] Figure 8 This is a schematic diagram of the resource scheduling process provided in an embodiment of this application;

[0054] Figure 9 This is a schematic block diagram of a communication device provided in an embodiment of this application;

[0055] Figure 10 This is another schematic block diagram of the communication device provided in the embodiments of this application;

[0056] Figure 11 This is a schematic diagram of a radio access network intelligent controller (RIC) architecture communication system shown in an embodiment of this application;

[0057] Figure 12 This is a schematic diagram of a communication system based on an open radio access network (ORAN or O-RAN) architecture provided in an embodiment of this application. Detailed Implementation

[0058] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0059] Before describing the technical solutions in this application, the following points should be noted.

[0060] First, in this application, the terms "first" and "second" are used to distinguish identical or similar items that have essentially the same function and purpose. For example, "first information" and "second information" are used merely to distinguish different information and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.

[0061] Second, in this application, the words "exemplarily" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design that is described as "exemplarily" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0062] Third, in this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0063] Fourth, in this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain instruction is called the information to be instructed. In specific implementation, there are many ways to instruct the information to be instructed, such as, but not limited to, directly instructing the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly instruct the information to be instructed by instructing other information, where there is a relationship between the other information and the information to be instructed; or it can only instruct a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing instruction overhead to some extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction, the instruction can be used to instruct the information to be instructed, and for the receiver of the instruction, the instruction can be used to determine the information to be instructed.

[0064] Fifth, in this application, "send" and "receive" indicate the direction of signal transmission. For example, "send information to a network device" can be understood as the destination of the information being the network device, which may include direct transmission via the air interface or indirect transmission via the air interface from other units or modules. "Receive information from a terminal device" can be understood as the source of the information being the terminal device, which may include direct reception from the terminal device via the air interface or indirect reception from the terminal device via the air interface from other units or modules. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface.

[0065] In other words, sending and receiving can occur between devices, such as between network devices and terminal devices; or they can occur within a device, such as between components, modules, chips, software modules, or hardware modules within a device via a bus, wiring, or interface.

[0066] It is understandable that information may undergo necessary processing, such as encoding and modulation, before being sent from the source to the destination. Similarly, the destination, upon receiving information from the source, can also perform corresponding processing, such as decoding and demodulation, to interpret the valid information from the source. Similar expressions in this application can be understood in a similar way and will not be elaborated further.

[0067] Sixth, the technical solutions of the embodiments of this application can be applied to various communication systems, such as: Long Term Evolution (LTE) systems, 5th Generation (5G) systems, or New Radio (NR) systems, or future communication systems, etc. This application does not limit them in this regard.

[0068] The following is a combination of... Figure 1 The communication system applicable to the embodiments of this application will be described in detail.

[0069] Figure 1 This is a schematic diagram of a communication system used in an embodiment of this application. Figure 1 A possible, non-limiting system schematic diagram is shown. For example... Figure 1 As shown, the communication system includes a RAN. Optionally, the communication system may also include a core network (CN) and the Internet.

[0070] The wireless access network may include at least one wireless access network device (which may be referred to as an access network device or a network device) and at least one terminal device. Figure 1This example uses one wireless access network (WLAN) device and two terminal devices. The terminal devices can connect to the WLAN device wirelessly. The WLAN device can connect to the core network wirelessly or via a wired connection. The core network device and the WLAN device can be different physical devices, or they can be the same physical device integrating core network and WLAN logical functions, or they can be the same physical device integrating some core network and some WLAN logical functions; this application does not limit this. Terminal devices can connect to each other, and WLAN devices can connect to each other, via wired or wireless connections.

[0071] Wireless access network (WLAN) devices and terminal devices can communicate via a wireless link. In one possible scenario, the WLAN device can act as a receiver, and the terminal device as a transmitter, sending signals to the WLAN device. However, this should not be construed as limiting this application. For example, in another possible scenario, the WLAN device can act as a transmitter, and the terminal device as a receiver, sending signals to the terminal device. Sidelink (SL) communication can also occur between terminal devices.

[0072] Optionally, the aforementioned radio access network can be a cellular system related to the 3rd Generation Partnership Project (3GPP), such as an LTE system, a 5G system, or a future communication system. The aforementioned radio access network can also be an O-RAN. The aforementioned radio access network can also be a cloud radio access network (CRAN), etc. This application does not limit this.

[0073] Understandable Figure 1 This application only illustrates one possible communication system architecture that can be applied to the embodiments of this application. In other possible scenarios, the above communication system may also include a greater number of wireless access network devices and terminals. The above communication system may also include other types of devices, such as relay devices and / or backhaul devices, which will not be listed here.

[0074] As an example and not a limitation, wireless access network devices can be used to help terminals achieve wireless access.

[0075] In one possible scenario, wireless access network equipment can be an evolved NodeB (eNodeB), a transmitting and receiving point (TRP), a micro base station, a transmitting point (TP), a next-generation NodeB (gNB), a base station in a future communication system, a high-altitude platform or satellite in a non-terrestrial network (NTN) communication system, or a wireless controller in a CRAN, etc.

[0076] In another possible scenario, multiple radio access network (RAN) devices collaborate to assist a terminal in achieving wireless access, with each RAN device performing a portion of the base station's functions. For example, RAN devices can be central units (CUs), distributed units (DUs), CU-control plane (CPs), CU-user plane (UPs), or radio units (RUs). CUs and DUs can be separate entities or included in the same network element; for example, CUs and DUs can be included in a baseband unit (BBU). RUs can be included in radio frequency (RF) devices or RF units; for example, RUs can be included in a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH). It is understood that RAN devices can be CU nodes, DU nodes, or devices comprising both CU and DU nodes.

[0077] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an open RAN system, CU can also be called O-CU (open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.

[0078] A terminal can also be called a terminal device, user equipment (UE), mobile station (MS), mobile terminal (MT), etc., or a device used to provide voice or data connectivity to users, or an Internet of Things (IoT) device. Currently, terminals can include, for example: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices (such as smartwatches, smart bracelets, pedometers, smart glasses, etc.), in-vehicle equipment (such as cars, bicycles, electric vehicles, airplanes, ships, trains, high-speed trains, etc.), satellite terminals, virtual reality (VR) devices, augmented reality (AR) devices, smart point-of-sale (POS) machines, customer-premises equipment (CPE), light user equipment (UE), reduced capability user equipment (REDCAPUE), wireless terminals in industrial control, smart home devices (such as refrigerators, televisions, air conditioners, electricity meters, etc.), smart robots, robotic arms, workshop equipment, wireless terminals in autonomous driving, wireless terminals in smart healthcare, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, and flying equipment (such as smart robots, hot air balloons, drones, airplanes), etc. Terminals can also be vehicle devices, such as vehicle units, vehicle modules, vehicle chips, on-board units (OBUs), or telematics boxes (T-BOXs).

[0079] To facilitate understanding of this application, some technical terms used in this application will be introduced below.

[0080] First, "large models": These can be understood as machine learning models with a large number of parameters and complex structures. These models can be trained on large-scale datasets to capture complex data patterns and features, thus performing well in various tasks. In contrast, they can be called "small models," that is, machine learning models with a small number of parameters and simple structures. For example, the teacher model described below can be considered a large model, and the student model can be considered a small model.

[0081] II. Knowledge Distillation: This is a model compression method that uses a teacher model to guide the training of a student model, thereby obtaining a lightweight student model. The loss function of the student model satisfies: loss = aL1 + (1-a)L2, where L1 represents the distillation loss, used to evaluate the similarity between the output of the student model and the output of the teacher model, L2 represents the task-specific loss, and 'a' represents a hyperparameter used to balance the distillation loss and the task-specific loss.

[0082] Figure 2 This is a schematic diagram of existing knowledge distillation.

[0083] like Figure 2 As shown, the teacher model receives input data and outputs inference results (or predictions), which include soft targets containing information about the similarity between categories. The student model receives input data and the aforementioned soft targets; that is, it is trained using the soft targets generated by the teacher model and the original input data. During training, the student model's loss function can be, for example, a weighted sum of the task's inherent loss and distillation loss. By optimizing the loss function, the student model's parameters are adjusted to improve its performance on the task, capturing and utilizing the knowledge from the teacher model.

[0084] III. Large Model Fine-Tuning Techniques: These techniques use personalized data to fine-tune large models, reducing storage, training, and inference overhead. For example, low-rank adaptation (LoRA) is a large model fine-tuning technique that reduces storage, training, and inference overhead by decomposing weight parameters into low-rank components. Another example is reinforcement learning from human feedback (RLHF), which involves: first, collecting human feedback, such as generating a series of outputs using a pre-trained language model, and having humans evaluate these outputs, such as selecting versions that align with human preferences. These evaluations are used to construct a dataset reflecting human preferences. Second, a reward model is trained using this human feedback to predict the probability that a particular output of the language model will be preferred by humans. Finally, the language model is fine-tuned.

[0085] IV. MCTS: A search model for decision-making processes, particularly suitable for game theory and planning problems. This model transforms decision problems into a tree-based node selection (or path selection) process. Each node represents a decision (or action), the parent node represents the previous decision, and the child nodes represent the next decision. Each node is obtained through expansion steps. When a node has expanded to all its child nodes, the path selection is completed based on the tree growth steps. The MCTS process mainly includes the following steps:

[0086] 1. Node expansion

[0087] If the selected node is not the terminating node (e.g., the game has not ended), one or more child nodes are generated from that node.

[0088] 2. Tree growth

[0089] Child nodes are selected using a strategy such as the upper confidence bound (UCB) algorithm. For example, this can be achieved by using Formula 1: Obtain the decision ak (child node), where Qa represents the environment score corresponding to decision a, pa represents the prior probability corresponding to decision a, Na represents the number of visits to decision a, ∑bNb represents the sum of the number of visits to all child nodes of this node, and c is a constant used to balance exploration and exploitation.

[0090] 3. Simulation

[0091] Starting with the newly expanded node, perform a complete simulation until the termination state is reached to obtain the environment score for that path.

[0092] 4. Backpropagation, also known as backtracking

[0093] Update the statistics (such as number of visits and environment score) for each node on the above path to optimize future selection strategies.

[0094] In the above process, the environment score refers to the reward or penalty given to the decision-maker by the environment after taking a certain action in a specific state. This environment score can be used to evaluate the merits of the action, enabling the model to make better choices in future decisions. For example, in the simulation phase of the MCTS process, starting from an expanded node, a complete simulation is executed until the terminal state is reached, obtaining the environment score for that path. For instance, in a board game, the environment score might represent winning the game (e.g., an environment score of 1), losing the game (e.g., an environment score of 0), or a draw (e.g., an environment score of 0.5). In path planning, the environment score might be related to the path length, time taken, or obstacle avoidance success rate. These are just a few examples.

[0095] Prior probability refers to the estimate of the probability of an event occurring before any observed data is available.

[0096] Figure 3 This is a schematic diagram of the MCTS process provided in the embodiments of this application.

[0097] like Figure 3As shown, taking node A in the third layer as an example, assuming node A is not a terminating node, then node A can generate one or more child nodes, such as... Figure 3 The example shows child nodes 5 and 6, each with a prior probability. For instance, child nodes 5 and 6, along with their corresponding prior probabilities, can be obtained based on a pre-trained model. It should be understood that child nodes 1 through 4, 7, and 8 are similar; for example, child nodes 7 and 8 can be child nodes corresponding to node B. The prior probabilities of child nodes 7 and 8, along with their corresponding prior probabilities, can also be obtained based on a pre-trained model, which will not be elaborated further here. Further, for node A, a child node can be selected based on Formula 1 above, such as child node 6. Then, starting from child node 6, a complete simulation is performed to obtain the environment score for that path (which includes child node 6, its parent node, and so on upwards to the root node). The process of selecting child node 5 is similar to that of child node 6, and will not be elaborated further here. Finally, the simulation results are passed upwards from child node 6 to the root node, updating the statistical information (such as the number of visits and the environment score) of each node on the path.

[0098] Currently, both in the fields of communications and computers, decision-making processes are frequently involved. With the development of user terminals, different users may have different needs; therefore, providing personalized decision-making is crucial. For example, a small model (such as a student model) can be deployed on the user terminal, guided by a teacher model, trained on the user terminal, and then used directly on the user terminal to make decisions. Another example is fine-tuning a large model using personalized data and deploying it on the user terminal to provide personalized decision-making. Yet another example is fine-tuning a large model on the network side and providing personalized decision-making services to the user terminal.

[0099] In personalized decision-making, optimizing the decision-making process is an important research area. Communication devices can often rely on predefined probabilities corresponding to each action to guide the decision-making process; for example, in the MCTS model, the decision-making process can be guided by predefined probabilities of each action. However, the accuracy of these predefined probabilities is relatively poor, thus affecting the effectiveness and reliability of the decisions made.

[0100] In view of this, this application provides a communication method in which a first device can deploy a reward model, obtain the reward corresponding to each action based on the reward model, and indicate the reward corresponding to the action to a network device so that the network device can update the prior probability corresponding to the action. By indicating the reward corresponding to the action by the first device, the network device can obtain which (or several) actions among the actions can obtain higher rewards, so that the network device can make better decisions. Moreover, the reward is more accurate than the predefined prior probability. Therefore, making decisions based on the updated probability is more effective and reliable than making decisions based on the prior probability corresponding to each action. In addition, the reward model deployed by the first device to obtain the reward of the action can be different for different users. This is conducive to realizing personalized decision-making of user terminals. For example, the first device takes a terminal device as an example. Thus, the reward model deployed or trained on different terminal devices can be different, so the rewards obtained may be different, which helps to realize personalized decision-making of user terminals.

[0101] The following is combined Figure 4 The differences between existing technologies and the methods provided in this application are given.

[0102] Figure 4 This is a comparative schematic diagram of existing technologies and methods provided in the embodiments of this application.

[0103] like Figure 4 As shown, in existing technologies, for a given node, if the node is not a terminating node, then the child nodes of that node and the prior probabilities of each child node are obtained based on a pre-trained model (e.g., ...). Figure 4 As shown in the diagram, the prior probability for action 0 is 0.4, and the prior probability for action 1 is 0.6. This means that the prior probability of each child node can be obtained based on a predefined model, or in other words, the prior probability of each child node is predefined, and decisions are made based on this prior probability. In this application, the reward for each child node can be obtained based on a reward model. The input to this reward model can be each child node, or each action. The aforementioned reward is used to update the prior probability of each child node (e.g., ...). Figure 4 As shown in the diagram, the prior probability for action 0 is updated to 0.2, and the prior probability for action 1 is updated to 0.8. The network device then makes decisions based on these updated probabilities, which facilitates better decision-making. Furthermore, the rewards mentioned above are more accurate than the prior probabilities. Therefore, making decisions based on the updated probabilities is more effective and reliable than making decisions based on the prior probabilities corresponding to each action.

[0104] The following is combined Figure 5This application provides a detailed description of the communication method. The embodiments shown in this application illustrate the method provided by this application from the perspective of communication device interaction. The specific forms and numbers of the communication devices shown are merely examples and should not constitute any limitation on the implementation of the method provided by this application. Below, taking a first terminal device and a first network device as the execution subjects, the communication method of the embodiments of this application will be described in detail. The first terminal device is an example of a first device, but it can also be replaced by DT, etc.

[0105] It should be understood that the first terminal device may be the first terminal device itself, or a chip, chip system or processor that supports the first terminal device in implementing the communication method, or a logic module or software that can implement all or part of the functions of the first terminal device; the first network device may be the first network device itself, or a chip, chip system or processor that supports the first network device in implementing the communication method, or a logic module or software that can implement all or part of the functions of the first network device, and this application does not make specific limitations in this regard.

[0106] Figure 5 This is a flowchart illustrating the communication method 500 provided in an embodiment of this application. This method 500 can, for example, be applied to... Figure 1 The system shown, the method 500 includes the following steps:

[0107] In step 510, the first terminal device determines the rewards corresponding to K actions based on the reward model.

[0108] Where K is a positive integer. The above action can be understood as a choice or operation that the device can perform in a certain state. The above action can also be called a decision, policy, choice, or operation, etc. This application does not limit its name, and this application does not exclude the possibility of using other names in future agreements.

[0109] The aforementioned reward can be understood as the feedback obtained by the device after performing a certain action, used to measure the quality of that action under a certain state. The aforementioned reward model can also be called an evaluation model, and the aforementioned reward can also be called a score, used to evaluate the quality of a certain action; this application does not limit its name.

[0110] It should be noted that the reward model described above can be pre-trained based on a dataset. The input to the reward model can be the K actions mentioned above, and the output can be the reward corresponding to each of the K actions. Each action can correspond to one reward; that is, there can be a one-to-one correspondence between actions and rewards. It can also be understood that the input to the reward model can include a state, and the K actions can be the K possible actions within that state.

[0111] For example, the first terminal device acquires the aforementioned K actions and inputs them into the aforementioned reward model to obtain K rewards corresponding to the aforementioned K actions. For instance, taking the MCTS model as an example, the aforementioned K actions can be K of one or more actions generated based on the MCTS model. More specifically, taking node A as an example, when node A is not a terminal node, node A can be expanded, such as generating one or more child nodes (an example of one or more actions mentioned above). The first terminal device can acquire K of the aforementioned one or more child nodes and input them into the aforementioned reward model to obtain K rewards corresponding to the aforementioned K child nodes. It can be understood that the input to the reward model can also include the state of node A, so that the reward model can output K rewards corresponding to the aforementioned K child nodes in the aforementioned states.

[0112] In one possible implementation, Figure 5 The method 500 also includes step 505, in which the first terminal device receives second information from the first network device. This second information is used to indicate the aforementioned K actions. In other words, the aforementioned K actions are indicated by the first network device, meaning that the first terminal device obtaining the K actions includes the first terminal device obtaining the aforementioned K actions from the first network device. That is, the first network device performs a decision-making process, indicating K actions to the first terminal device, or in other words, indicating K choices / decisions. The first terminal device determines the reward for the K actions based on a reward model. Optionally, the first network device may also indicate a state to the first terminal device, and the aforementioned K actions may be K possible actions in that state.

[0113] For example, taking the MCTS model as an example, the first network device can deploy the MCTS model. During the node expansion process, taking the expansion of node A as an example, the first network device obtains one or more child nodes (i.e., one or more actions) of node A and the prior probabilities corresponding to each child node, and indicates the K child nodes (i.e., K actions) with the highest prior probabilities among the one or more child nodes to the first terminal device. The prior probabilities corresponding to the one or more child nodes can be predefined, for example, a pre-defined pre-trained model. The one or more child nodes and their corresponding prior probabilities can be obtained by the first network device based on the pre-trained model, such as inputting the actions and states of node A into the pre-trained model, and the pre-trained model outputting the one or more child nodes and their corresponding prior probabilities.

[0114] By selecting the K child nodes with the highest prior probabilities, it is beneficial to balance performance and complexity. For example, selecting K child nodes can reduce computational complexity, while selecting the K with the highest prior probabilities allows the device to concentrate resources and computing power on actions that are most likely to bring high rewards, which helps to improve the efficiency and effectiveness of decision-making.

[0115] It is understandable that if one or more of the above child nodes do not have corresponding prior probabilities, such as in a scenario where the first network device does not need a pre-trained model, the first network device can randomly select K child nodes to indicate to the first terminal device.

[0116] In step 520, the first terminal device sends first information to the first network device, which is used to indicate the reward corresponding to the above K actions.

[0117] The first information is used to update the prior probabilities corresponding to the aforementioned K actions. After determining the rewards corresponding to the aforementioned K actions, the first terminal device can indicate the rewards corresponding to the aforementioned K actions to the first network device, so that the first network device can update the prior probabilities corresponding to the aforementioned K actions.

[0118] The first information mentioned above can be designed in two ways: one is that it includes K rewards corresponding to the K actions; the other is that it includes K probabilities corresponding to the K actions, where the K probabilities are normalized from the K rewards. These K probabilities can represent the expected good or bad of the K actions. In one possible implementation, the action with the higher probability is more likely to be selected.

[0119] Normalizing the above K rewards, for example, can be done by transforming each reward through a function (such as the "softmax" function) to ensure that the above K rewards are all positive numbers. Then, the transformed reward value is divided by their sum to obtain the probability corresponding to each of the above K actions.

[0120] In step 530, the first network device updates the prior probabilities corresponding to the K actions based on the first information.

[0121] In one possible implementation, if the first information includes the K rewards corresponding to the K actions, the first network device normalizes the K rewards to obtain K probabilities, and updates the prior probabilities corresponding to the K actions based on these K probabilities. The method for normalizing the K rewards can be found in step 520, and will not be detailed here.

[0122] In another possible implementation, if the first information includes the K probabilities corresponding to the K actions, the first network device updates the prior probabilities corresponding to the K actions based on the K probabilities.

[0123] It's easy to understand that after updating the prior probabilities, the probabilities corresponding to high-reward actions may be higher, while the probabilities corresponding to low-reward actions may be lower. By optimizing the probabilities of actions, when the first network device makes a decision / selection based on the probabilities of the aforementioned K actions, the reliability and effectiveness of the decision are higher. For the model, this leads to better performance.

[0124] After updating the prior probabilities corresponding to the above K actions, the first network device can make a decision from the above K actions based on the updated probabilities, such as selecting one of the above K actions.

[0125] For example, when the above method is applied to the MCTS model, after the first network device updates the prior probabilities corresponding to the above K actions, it can be done using Formula 2: Obtain action ak, where qa represents the score corresponding to action a, Pa represents the updated prior probability corresponding to action a, Na represents the number of visits to action a, ∑bNb represents the sum of the number of visits to all child nodes of this node, and c is a constant used to balance exploration and exploitation.

[0126] The score corresponding to action 'a' can be a weighted average of the environment score of action 'a' and the reward for action 'a' obtained based on the reward model. For example, the score 'qa' corresponding to action 'a' satisfies 'qa = m·n + (1-m)·l', where n represents the environment score of action 'a', l represents the reward for action 'a', and m represents the weight.

[0127] It should be noted that the above l can be the reward for action a obtained based on the reward model, or it can be replaced by the normalized probability of the reward for action a obtained based on the reward model. When l is the reward for action a obtained based on the reward model, the corresponding weights can be different from when l is the normalized probability of the reward for action a obtained based on the reward model.

[0128] In one possible implementation, the method 500 further includes steps 535 and 540. In step 535, the first network device sends third information to the first terminal device, the third information indicating a first path; correspondingly, the first terminal device receives the third information. In step 540, the first terminal device sends a score corresponding to the first path to the first network device. Correspondingly, the first network device receives the score of the first path.

[0129] The aforementioned first path can be understood as the set of decisions made by the first network device. For example, in the MCTS model, the first path can be understood as the set of nodes selected by each layer of the tree. The aforementioned first path can also be called a decision set, decision sequence, action set, action sequence, etc., which will not be listed here, and this application does not limit its name.

[0130] For example, refer to Figure 3 As shown in the example, the first path may include child node 6, the parent node of child node 6, and so on, up to the root node. The first network device can indicate the first path to the first terminal device to obtain the score of the first path, thereby facilitating the backtracking of the statistical information corresponding to each node on the first path, such as the number of visits and the score.

[0131] In the above scheme, the first network device indicates the first path to the first terminal device to obtain the score of the first path. In this way, on the one hand, the computational burden of the first network device can be reduced, and on the other hand, since different user terminals have different environments, channels, etc., their feedback scores of the first path may also be different. Thus, the first network device can obtain personalized feedback for different terminal devices.

[0132] In one possible implementation, the above score is a weighted average of the environment score of the first path and the reward for the first path obtained based on the reward model. For example, the score r corresponding to the first path satisfies r = x·v + (1-x)·g, where v represents the environment score of the first path, g represents the reward for the first path, and x represents the weight.

[0133] It should be noted that the above g can be the reward of the first path obtained based on the reward model, or it can be replaced by the probability after normalization of the reward of the first path obtained based on the reward model. When g is the reward of the first path obtained based on the reward model, the corresponding weights can be different from when g is the probability after normalization of the reward of the first path obtained based on the reward model.

[0134] In the above scheme, by comprehensively considering the environmental score and the reward obtained based on the reward model, a more comprehensive and accurate evaluation can be obtained, thereby improving the reliability and accuracy of the decision-making process.

[0135] Optionally, in Figure 5 In the method shown, prior to step 505, the method further includes: the first network device initiating a decision or the first terminal device initiating a decision. For example, in a scenario where the first network device needs to make a decision, the first network device may initiate the decision, such as in a scenario where the first network device performs beam prediction or resource scheduling. As another example, in a scenario where the first terminal device needs to make a decision, the first terminal device may send the decision.

[0136] In one possible implementation, the method 500 further includes step 545, in which the first network device sends fourth information to indicate K first probabilities corresponding to the K actions. These K first probabilities are updated prior probabilities associated with the K actions by the first terminal device, or they are obtained based on updated prior probabilities associated with multiple terminal devices, including the first terminal device, each of which is associated with updated prior probabilities corresponding to the K actions. Accordingly, the first terminal device receives the fourth information. Further, the first terminal device can indicate the K first probabilities corresponding to the K actions to the second network device.

[0137] In other words, the first network device can send the first probabilities corresponding to the aforementioned K actions to the first terminal device. This allows another network device, such as the second network device, to directly report the first probabilities corresponding to the aforementioned K actions to the second network device during decision-making. The second network device does not need to request rewards corresponding to the aforementioned K actions from the first terminal device to update its stored prior probabilities. This not only reduces the computational load on the first terminal device but also reduces signaling overhead.

[0138] One possible design for the aforementioned K first probabilities is that these K first probabilities are the updated prior probabilities corresponding to the K actions associated with the first terminal device. For example, after the first network device updates the prior probabilities corresponding to the K actions, it sends the K updated prior probabilities to the first terminal device.

[0139] Another possible design is that the K first probabilities are obtained based on updated prior probabilities associated with multiple terminal devices, including a first terminal device. Each terminal device is associated with K actions corresponding to K updated prior probabilities. For example, the K first probabilities are obtained by fusing the updated prior probabilities associated with multiple terminal devices. For instance, the K first probabilities are obtained by averaging the updated prior probabilities associated with multiple terminal devices. For example, if the multiple terminal devices include terminal device 1 and terminal device 2, the K updated prior probabilities associated with terminal device 1 for the K actions include: P1 (prior probability of action 1), P2 (prior probability of action 2), and P3 (prior probability of action 3). The K updated prior probabilities associated with terminal device 2 for the K actions include: P4 (prior probability of action 1), P5 (prior probability of action 2), and P6 (prior probability of action 3). Then the K first probabilities are (P1+P4) / 2, (P2+P5) / 2, and (P3+P6) / 2. For example, the aforementioned K first probabilities are obtained by weighted averaging of the updated prior probabilities associated with multiple terminal devices. Or, the aforementioned K first probabilities are obtained by taking the maximum / minimum value of the updated prior probabilities associated with multiple terminal devices. This application does not limit the specific probabilities in this regard.

[0140] Figure 6 This is a schematic diagram illustrating the interaction between the first terminal device and the second network device provided in an embodiment of this application. Figure 6 One possible example of interaction between a first terminal device and a second network device is given, but this should not be construed as limiting the application in any way. There may be more or fewer interaction steps between the first terminal device and the second network device.

[0141] In step 610, the first terminal device and the second network device perform model alignment.

[0142] For example, the first terminal device may send a first identifier to the second network device. This first identifier identifies a reward model deployed on the first terminal device. The first identifier can be a reward model identifier or a probability identifier. It is understood that the probability identifier can identify the reward model corresponding to the following K first probabilities. The second network device receives the first identifier. The second network device may send auxiliary information for the computation task to the first terminal device. This auxiliary information may include, for example, the configuration of the computation task. For instance, in a beam prediction scenario, the second network device may indicate to the first terminal device one or more candidate beams to be predicted and / or the number of such candidate beams.

[0143] It should be understood that the alignment of the above models is merely an example and should not constitute any limitation on this application. For example, it can also be applied to the alignment of functions.

[0144] In step 620, the first terminal device sends K first probabilities to the second network device.

[0145] After the first terminal device and the second network device perform model alignment, the first terminal device reports the K first probabilities corresponding to the aforementioned K actions to the second network device.

[0146] Optionally, Figure 6 The method also includes steps 605 and 606. In step 605, the second network device sends a capability request to the first terminal device to request capability information of the first terminal device. The first terminal device receives the capability request. In step 606, the first terminal device sends capability information to the second network device. This application does not limit the content of the capability information.

[0147] It should be noted that, Figure 5 The method described uses the first path as an example; for any path, it can be followed... Figure 5 The method shown is executed. In one possible implementation, a stopping condition can be predefined, and the first network device repeats the decision-making process multiple times, one path at a time, until the stopping condition is met.

[0148] The following will combine Figure 7 and Figure 8 Give Figure 5 The method 500 shown is a detailed process for beam prediction scenarios and a detailed process for resource scheduling scenarios.

[0149] Figure 7 This is a schematic diagram of the beam prediction process provided in the embodiments of this application.

[0150] In step 701, channel measurement and reporting are performed between terminal device 1 (an example of a first terminal device) and network device 1 (an example of a first network device).

[0151] One possible example is that, based on downlink signals, channel measurements are performed. Network device 1 transmits a downlink reference signal, and terminal device 1 receives the downlink reference signal and performs channel measurements to obtain measurement results. Terminal device 1 can report the measurement results to network device 1. The measurement results may include, but are not limited to, the identifier of the beam corresponding to the downlink reference signal and the beam quality. The beam quality can be characterized by at least one of the following parameters: received signal strength indicator (RSSI), reference signal received power (RSRP), reference signal receiving quality (RSRQ), signal-to-noise ratio (SNR), or signal-to-interference plus noise ratio (SINR). Beam quality can also be characterized by other parameters, which are not limited in this application.

[0152] The aforementioned downlink reference signals include, but are not limited to: channel state information reference signal (CSI-RS), cell specific reference signal (CS-RS), user equipment specific reference signal (US-RS), demodulation reference signal (DMRS), or synchronization signal block (SSB), etc.

[0153] Network device 1 can use the above measurement results as the initial value for beam prediction. For example, the initial value for beam prediction can be the beam identifier, RSRP, RSRQ, or RSSI, etc., which will not be listed here.

[0154] In step 702, network device 1 performs beam prediction.

[0155] For example, network device 1 can perform beam prediction based on the pre-trained model and the initial values ​​of the beam prediction described above to obtain one or more candidate beams. This process can be viewed as the process of generating one or more child nodes during node expansion in the MCTS model.

[0156] In step 703, network device 1 sends a reference signal resource configuration to terminal device 1, and correspondingly, terminal device 1 receives the reference signal resource configuration.

[0157] The aforementioned reference signal resource configuration is used to configure the measurement resources of the reference signals corresponding to the aforementioned K candidate beams, so that the terminal device 1 can measure the quality of the corresponding candidate beams based on the measurement resources. The aforementioned K candidate beams are K beams selected from the aforementioned one or more candidate beams. For example, the aforementioned K candidate beams can be K beams randomly selected from the aforementioned one or more candidate beams, or K beams with the best signal quality (e.g., maximum RSRP) from the aforementioned one or more candidate beams, or K beams with the highest prior probability from the aforementioned one or more candidate beams. The prior probability corresponding to each candidate beam can be obtained based on a pre-trained model; that is, in step 702, the output of the pre-trained model can be one or more candidate beams and the prior probability corresponding to each candidate beam.

[0158] For example, network device 1 sends CSI resource configuration to terminal device 1 to configure the CSI resources corresponding to the above K candidate beams, and each CSI resource is used to measure the quality of the corresponding candidate beam.

[0159] In step 704, terminal device 1 performs measurements based on the reference signal resource configuration.

[0160] For example, terminal device 1 measures the quality of the above K candidate beams based on CSI resource configuration and obtains the measurement results corresponding to the K candidate beams.

[0161] In step 705, terminal device 1 determines the reward corresponding to the above K candidate beams based on the reward model.

[0162] Among them, the above K candidate beams can be regarded as Figure 5 The method shown is an example of K actions in 500.

[0163] In one possible implementation, the reward corresponding to the K candidate beams is obtained based on the reward model and the measurement results corresponding to the K candidate beams. For example, the input to the reward model may include, but is not limited to, the measurement results corresponding to the K candidate beams. The input may also include auxiliary information, such as the moving speed of terminal device 1, environmental information around the location of terminal device 1, and interference information experienced by terminal device 1.

[0164] In step 706, terminal device 1 sends a CSI report to network device 1. Accordingly, network device 1 receives the aforementioned CSI report.

[0165] In addition to the measurement results of the reference signal, the CSI report mentioned above can also indicate the K rewards corresponding to the K candidate beams, such as carrying the K rewards corresponding to the K candidate beams or the K probabilities after normalization of the K rewards.

[0166] In step 707, network device 1 updates the prior probabilities corresponding to the above K candidate beams.

[0167] Based on the K rewards / K probabilities corresponding to the K candidate beams reported by the terminal device 1, network device 1 updates the prior probabilities corresponding to the K candidate beams.

[0168] In step 708, network device 1 sends the updated prior probabilities corresponding to the K candidate beams to terminal device 1. Correspondingly, terminal device 1 receives the updated prior probabilities corresponding to the K candidate beams. This facilitates terminal device 1's ability to directly report the updated prior probabilities corresponding to the K candidate beams to other network devices subsequently.

[0169] Step 708 can be optional. In some implementations, network device 1 may not send the updated prior probabilities corresponding to the above K candidate beams to terminal device 1.

[0170] Understandable. Figure 7 This paper takes one terminal device as an example, but this should not be construed as limiting this application. More terminal devices can also evaluate the above K candidate beams based on their respective deployed reward models. The process is the same as that of terminal device 1, and will not be repeated here.

[0171] When network device 1 receives K rewards corresponding to K candidate beams associated with multiple terminal devices, network device 1 can update the prior probabilities corresponding to the K candidate beams associated with each terminal device, and based on the K updated prior probabilities, obtain K first probabilities and indicate them to terminal device 1. The specific method for obtaining the first probabilities can be found in [reference needed]. Figure 5 The relevant explanations will not be repeated here.

[0172] It can also be understood that, during the backtracking phase, the updated statistical information for each node can be its score and the number of visits. This score can be a weighted average of the environment score and the reward obtained based on the reward model. The aforementioned score can be the accuracy of beam prediction and the mean square error of the predicted beam quality. Beam quality can be characterized by parameters such as RSRP, RSRQ, and SINR. Specific characterizable parameters can be found in step 701.

[0173] exist Figure 7The method shown improves the accuracy and reliability of beam prediction by updating the prior probabilities corresponding to the predicted candidate beams.

[0174] Figure 8 This is a schematic diagram of the resource scheduling process provided in the embodiments of this application.

[0175] In step 801, channel measurement and reporting are performed between terminal device 1 and network device 1 (an example of the first network device).

[0176] One possible example is channel measurement based on downlink signals. Network device 1 sends a downlink reference signal, terminal device 1 receives the downlink reference signal, performs channel measurement, and obtains the measurement results. Terminal device 1 can report the measurement results to network device 1. The measurement results may include, but are not limited to, the identifier of the beam corresponding to the downlink reference signal and beam quality. Parameters that can be used to characterize beam quality can be found in [reference needed]. Figure 7 This will not be elaborated upon here.

[0177] In step 802, terminal device 1 sends a status to DT. DT is an example of a first device.

[0178] The Data Transmission (DT) can create a virtual model of a physical entity or system for simulation, analysis, and optimization. The DT can be deployed in network devices or other locations, without limitation in this application. The DT is an example of the first device. The aforementioned states include, but are not limited to, one or more of the following: the size of the buffered packets, the buffered packet duration, the average rate, the user channel, the current rate, etc.

[0179] In step 803, network device 1 predicts one or more candidate scheduled users.

[0180] For example, the scheduler in network device 1 can predict one or more candidate scheduling users based on the above measurement results.

[0181] In step 804, network device 1 sends indication information 1 to DT, which indicates K candidate scheduling users. Accordingly, DT receives the aforementioned indication information 1.

[0182] The aforementioned K candidate scheduling users can be K of the one or more candidate scheduling users mentioned above. The aforementioned K candidate scheduling users are an example of the aforementioned K actions.

[0183] In step 805, DT determines the rewards corresponding to the above K candidate scheduling users based on the reward model.

[0184] By virtually scheduling in the DT (Distribution Technology), one or more of the following can be obtained as a reward: system throughput, fairness, or packet loss rate; or a weighted average of these parameters can be used as a reward. The input to the reward model can include one or more of the following: K candidate scheduling users, or the status of candidate scheduling users in the DT, such as the size of the buffered packets, the buffered packet time, the average rate, the channel of the candidate scheduling user, etc.

[0185] In step 806, DT sends indication information 2 to network device 1, which indicates the reward corresponding to the aforementioned K candidate scheduling users. Correspondingly, network device 1 receives the indication information 2. The indication information 2 is an example of the first information.

[0186] In step 807, network device 1 updates the prior probabilities corresponding to the above K candidate scheduling users based on indication information 2.

[0187] In step 808, network device 1 sends the updated prior probabilities corresponding to the K candidate scheduling users to DT. Correspondingly, DT receives the updated prior probabilities corresponding to the K candidate scheduling users. This facilitates DT's ability to directly report the updated prior probabilities corresponding to the K candidate scheduling users to other network devices subsequently. A detailed explanation of the updated prior probabilities corresponding to the K candidate scheduling users can be found in step 708, and will not be elaborated upon here.

[0188] exist Figure 8 The method shown improves scheduling performance and reliability by updating the prior probabilities of the predicted candidate scheduling users.

[0189] It should be noted that the order of the methods listed above does not imply the order of execution. The execution order of each process should be determined by its function and internal logic.

[0190] The communication method of the embodiments of this application has been described in detail above. The communication device of the embodiments of this application will be described in detail below. The communication device includes modules or units for performing each part of the above embodiments. The modules or units may be software, hardware, or a combination of software and hardware. The following is only a brief illustrative example of the communication device. For details of the implementation, please refer to the description of the foregoing method embodiments, which will not be repeated below.

[0191] Figure 9 This is a schematic block diagram of a communication device provided in an embodiment of this application. For example... Figure 9 As shown, the communication device 900 includes a processing module 910 and a transceiver module 920.

[0192] In one possible implementation, the communication device 900 is used to implement the steps corresponding to the first terminal device (an example of the first device) in the method 500 described above.

[0193] The processing module 910 is used to determine the rewards corresponding to K actions based on the reward model, where K is a positive integer; the transceiver module 920 is used to send first information to the first network device, which is used to indicate the rewards corresponding to the K actions and to update the prior probabilities corresponding to the K actions.

[0194] Optionally, the transceiver module 920 is further configured to receive second information from the first network device, the second information being used to indicate the aforementioned K actions.

[0195] Optionally, the above K actions are K of one or more actions generated based on the MCTS model.

[0196] Optionally, the transceiver module 920 is further configured to receive third information from the first network device, the third information being used to indicate the first path; and to send the score corresponding to the first path to the first network device.

[0197] Optionally, the above score is a weighted average of the environment score of the first path and the reward of the first path obtained based on the reward model.

[0198] Optionally, when the above method is applied to the first network device for beam prediction, the above K actions include the K candidate beams predicted by the first network device.

[0199] Optionally, the rewards corresponding to the above K actions are obtained based on the reward model and the measurement results corresponding to the K candidate beams.

[0200] Optionally, the aforementioned first information is carried in CSI.

[0201] Optionally, when the above method is applied to the first network device for resource scheduling, the above K actions include the K candidate scheduling users predicted by the first network device.

[0202] Optionally, the transceiver module 920 is further configured to receive fourth information from the first network device, the fourth information being used to indicate K first probabilities corresponding to the K actions, the K first probabilities being updated prior probabilities corresponding to the K actions associated with the first terminal, or the K first probabilities being obtained based on updated prior probabilities associated with multiple terminals, the multiple terminals including the first terminal, each of the multiple terminals being associated with updated prior probabilities corresponding to the K actions; and to indicate the K first probabilities to the second network device.

[0203] In one possible implementation, the communication device 900 is used to implement the steps corresponding to the first network device in the method 500 described above.

[0204] The transceiver module 920 is used to receive first information from the first device, which indicates the reward corresponding to K actions. The reward corresponding to the K actions is obtained based on a reward model, where K is a positive integer. The processing module 910 is used to update the prior probabilities corresponding to the K actions based on the first information.

[0205] Optionally, the transceiver module 920 is further configured to send second information to the first device, the second information being used to indicate the K actions.

[0206] Optionally, the above K actions are K of one or more actions generated based on the MCTS model.

[0207] Optionally, the transceiver module 920 is further configured to send third information to the first device, the third information being used to indicate the first path; and to receive the score corresponding to the first path.

[0208] Optionally, the above score is a weighted average of the environment score of the first path and the reward of the first path obtained based on the reward model.

[0209] Optionally, when the above method is applied to the first network device for beam prediction, the above K actions include the K candidate beams predicted by the first network device.

[0210] Optionally, the rewards corresponding to the above K actions are obtained based on the reward model and the measurement results corresponding to the K candidate beams.

[0211] Optionally, the aforementioned first information is carried in CSI.

[0212] Optionally, when the above method is applied to the first network device for resource scheduling, the above K actions include the K candidate scheduling users predicted by the first network device.

[0213] Optionally, the transceiver module 920 is further configured to send fourth information to the first device, the fourth information being used to indicate K first probabilities corresponding to the K actions, the K first probabilities being updated prior probabilities corresponding to the K actions associated with the first terminal, or the K first probabilities being obtained based on updated prior probabilities associated with multiple terminals, the multiple terminals including the first terminal, each of the multiple terminals being associated with updated prior probabilities corresponding to the K actions.

[0214] It should be understood that the communication device 900 here is embodied in the form of a functional module. The term "module" here can refer to application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors, etc.) and memories for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the communication device 900 can specifically be the first terminal device, the first network device, or the DT in the above embodiments. The communication device 900 can be used to execute the various processes and / or steps corresponding to the first terminal device, the first network device, or the DT in the above method embodiments; to avoid repetition, these will not be described further here.

[0215] The aforementioned communication device 900 has the function of implementing the corresponding steps performed by the first device (such as the first terminal device, terminal device 1, DT) or network device (such as the first network device, the second network device) in the aforementioned method; the aforementioned functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions. In the embodiments of this application, Figure 9 The communication device 900 in the text can also be a chip, such as a SoC.

[0216] It should be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division. In actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0217] Figure 10 This is another schematic block diagram of the communication device provided in the embodiments of this application.

[0218] The communication device 1000 can be a chip system, or an apparatus configured with a chip system to implement the methods described in the above-described method embodiments. In the embodiments of this application, the chip system can be composed of chips, or it can include chips and other discrete devices.

[0219] like Figure 10 As shown, the communication device 1000 may include a processor 1010, which can be used to execute computer programs or instructions stored in memory to achieve... Figures 5 to 8The steps performed by the first device (such as the first terminal device, terminal device 1, DT) or the network device (such as the first network device, the second network device) in any of the embodiments shown.

[0220] The communication device 1000 also includes a communication interface 1020. The communication interface 1020 can be used to communicate with other devices via a transmission medium, thereby enabling the communication device 1000 to communicate with other devices. The communication interface 1020 can be, for example, a transceiver, interface, pin, bus, circuit, or a device capable of transmitting and receiving functions. The processor 1010 can use the communication interface 1020 to input and output data and to implement... Figures 5 to 8 The steps performed by the first device (such as the first terminal device, terminal device 1, DT) or the network device (such as the first network device, the second network device) in any of the embodiments shown.

[0221] In one possible implementation, the communication device 1000 further includes at least one memory 1030 for storing program instructions and / or data. The memory 1030 is coupled to the processor 1010. The coupling in this embodiment is an indirect coupling or communication connection between devices, units, or modules, and can be electrical, mechanical, or other forms, used for information exchange between devices, units, or modules. The processor 1010 may operate in conjunction with the memory 1030. The processor 1010 may execute program instructions stored in the memory 1030. At least one of the at least one memory may be included in the processor.

[0222] It should be understood that the coupling in the embodiments of this application is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, used for information interaction between devices, units, or modules. The processor 1010 may operate in conjunction with the memory 1030. The embodiments of this application do not limit the specific connection medium between the processor 1010, communication interface 1020, and memory 1030. Optionally, the processor 1010, communication interface 1020, and memory 1030 are connected via a bus 1040. The bus 1040 is in... Figure 10 The connections between other components are shown in bold lines only and are not intended to be limiting. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0223] Figure 11 This is a schematic diagram of a RAN intelligent controller (RIC) architecture communication system illustrated in an embodiment of this application. Figure 11 As shown, the RIC architecture communication system includes an RIC. This RIC includes near-real-time RIC (near-RT RIC) and non-real-time RIC (non-RT RIC). Near-real-time RIC can be used for model training and inference, for example, to train artificial intelligence (AI) models and perform inference using those AI models. Near-real-time RIC can obtain network-side and / or terminal-side information from RAN devices (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data. Optionally, near-real-time RIC can send inference results to RAN devices and / or terminals. Optionally, inference results can be exchanged between CU and DU, and / or between DU and RU. For example, near-real-time RIC delivers inference results to DU, and DU sends them to RU, etc.

[0224] Non-real-time RICs can also be used for model training and inference. For example, they can be used to train AI models and then use those models for inference. Non-real-time RICs can obtain network-side and / or terminal-side information from RAN devices (e.g., CUs, CU-CPs, CU-UPs, DUs, and / or RUs) and / or terminals. This information can be used as training data or inference data, and the inference results can be delivered to the RAN devices and / or terminals. Optionally, inference results can be exchanged between CUs and DUs, and / or between DUs and RUs; for example, a non-real-time RIC delivers the inference result to a DU, which then forwards it to an RU.

[0225] Near real-time RICs and non-real-time RICs can also be configured as separate network elements. Optionally, near real-time RICs and non-real-time RICs can also be part of other devices. For example, near real-time RICs can be set in RAN nodes (e.g., CUs or DUs), while non-real-time RICs can be set in Operations Administration and Maintenance (OAM), cloud servers, core network devices, or other network devices.

[0226] It should be understood that Figure 11 The system shown can be applied to an open RAN architecture, for example. The CU can also be called O-CU, the DU can also be called O-DU, and the RU can also be called O-RU. They will not be listed here.

[0227] Figure 12 This is a schematic diagram of a communication system based on an open RAN architecture provided in an embodiment of this application. Figure 12 The open RAN architecture shown is merely an example; it can include various other architectures. Figure 12 Other components besides those shown. It should be understood that, Figure 12 CU can also be called O-CU, DU can also be called O-DU, and RU can also be called O-RU. They will not be listed here.

[0228] In this communication system, network elements are connected via interfaces (e.g., NG, Xn) or air interfaces (Uu). These network element nodes, such as core network equipment, access network equipment, and terminal equipment, may each have one or more AI modules (only one is shown in the figure for clarity). Access network equipment can be a single RAN device or can include multiple RAN devices, for example, CU and DU. The CU and / or DU may also each have one or more AI modules. Optionally, the CU can be further divided into CU-CP and CU-UP. One or more AI modules may each be configured in the CU-CP and / or CU-UP.

[0229] The AI ​​modules described above can be used to implement corresponding AI functions. AI modules deployed in different network elements can be the same or different. Depending on the parameter configuration, each AI module can perform different functions. An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or they can be deployed on the same node or device.

[0230] For example, in this application, step 520 can be implemented as follows: the DU corresponding to the first network device receives the first information through the RU. In the O-RAN system, step 520 can be implemented as follows: the O-DU corresponding to the first network device receives the first information through the O-RU. Step 530 can be implemented as follows: the DU / CU / RU corresponding to the first network device updates the prior probabilities corresponding to K actions based on the first information. In the O-RAN system, step 530 can be implemented as follows: the O-DU / O-CU / O-RU / RIC corresponding to the first network device updates the prior probabilities corresponding to K actions based on the first information.

[0231] This application also provides a computer program product, which includes: a computer program (also referred to as code or instructions), which, when run, can achieve... Figures 5 to 8The method described in any of the embodiments shown.

[0232] This application also provides a computer-readable storage medium storing a computer program (also referred to as code or instructions). When the computer program is executed, it can achieve... Figures 5 to 8 The method described in any of the embodiments shown.

[0233] This application provides a communication system, which includes a first terminal device and a first network device as described above.

[0234] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by the integrated logic circuitry in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0235] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0236] The terms "unit," "module," etc., used in this specification can be used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. In the embodiments of this application, "unit" and "module" have the same meaning and can be used interchangeably.

[0237] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the several embodiments provided in this application, it should be understood that the disclosed apparatus, devices, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0238] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0239] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0240] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs, DVDs), or semiconductor media (e.g., solid-state disks, SSDs), etc.

[0241] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the technology, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0242] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A communication method, characterized in that, Applied to a first device, the method includes: Based on the reward model, determine the rewards for K actions, where K is a positive integer; Send first information to the first network device. The first information is used to indicate the reward corresponding to the K actions and to update the prior probability corresponding to the K actions.

2. The method as described in claim 1, characterized in that, The method further includes: Receive second information from the first network device, the second information being used to indicate the K actions.

3. The method as described in claim 1 or 2, characterized in that, The method further includes: Receive third information from the first network device, the third information being used to indicate the first path; Send the score corresponding to the first path to the first network device.

4. The method as described in claim 3, characterized in that, The score is a weighted average of the environment score of the first path and the reward for the first path obtained based on the reward model.

5. The method according to any one of claims 1 to 4, characterized in that, When the method is applied to the first network device for beam prediction, the K actions include the K candidate beams predicted by the first network device.

6. The method as described in claim 5, characterized in that, The rewards corresponding to the K actions are obtained based on the reward model and the measurement results corresponding to the K candidate beams.

7. The method as described in claim 5 or 6, characterized in that, The first information is carried in the Channel State Information (CSI).

8. The method according to any one of claims 1 to 4, characterized in that, When the method is applied to the first network device for resource scheduling, the K actions include K candidate scheduling users predicted by the first network device.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: The system receives fourth information from the first network device. The fourth information is used to indicate K first probabilities corresponding to the K actions. The K first probabilities are updated prior probabilities corresponding to the K actions associated with the first terminal, or the K first probabilities are obtained based on updated prior probabilities associated with multiple terminals, including the first terminal. Each of the multiple terminals is associated with the updated prior probability corresponding to the K actions. Indicate the K first probabilities to the second network device.

10. The method according to any one of claims 1 to 9, characterized in that, The K actions are K of one or more actions generated based on the Monte Carlo Tree Search (MCTS) model.

11. A communication method, characterized in that, Applied to a first network device, the method includes: Receive first information from a first device, the first information being used to indicate the rewards corresponding to K actions, the rewards corresponding to the K actions being obtained based on a reward model, where K is a positive integer; Based on the first information, update the prior probabilities corresponding to the K actions.

12. The method as described in claim 11, characterized in that, The method further includes: Send a second message to the first device, the second message being used to instruct the K actions.

13. The method as described in claim 11 or 12, characterized in that, The method further includes: Send third information to the first device, the third information being used to indicate the first path; Receive the score corresponding to the first path.

14. The method as described in claim 13, characterized in that, The score is a weighted average of the environment score of the first path and the reward for the first path obtained based on the reward model.

15. The method according to any one of claims 11 to 14, characterized in that, When the method is applied to the first network device for beam prediction, the K actions include the K candidate beams predicted by the first network device.

16. The method as described in claim 15, characterized in that, The rewards corresponding to the K actions are obtained based on the reward model and the measurement results corresponding to the K candidate beams.

17. The method as described in claim 15 or 16, characterized in that, The first information is carried in the Channel State Information (CSI).

18. The method according to any one of claims 11 to 14, characterized in that, When the method is applied to the first network device for resource scheduling, the K actions include K candidate scheduling users predicted by the first network device.

19. The method according to any one of claims 11 to 18, characterized in that, The method further includes: Send a fourth message to the first device, the fourth message being used to indicate K first probabilities corresponding to the K actions, the K first probabilities being updated prior probabilities corresponding to the K actions associated with the first terminal, or the K first probabilities being obtained based on updated prior probabilities associated with multiple terminals, the multiple terminals including the first terminal, each of the multiple terminals being associated with the updated prior probabilities corresponding to the K actions.

20. The method according to any one of claims 11 to 19, characterized in that, The K actions are K of one or more actions generated based on the Monte Carlo Tree Search (MCTS) model.

21. A communication device, characterized in that, It includes modules for implementing the method as described in any one of claims 1 to 10, or includes modules for implementing the method as described in any one of claims 11 to 20.

22. A communication device, characterized in that, Includes a processor for invoking a computer program in memory to cause the communication device to implement the method as described in any one of claims 1 to 10, or to implement the method as described in any one of claims 11 to 20.

23. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 10, or implements the method as described in any one of claims 11 to 20.

24. A computer program product, characterized in that, The computer program product includes instructions that, when executed, implement the method as described in any one of claims 1 to 10, or implement the method as described in any one of claims 11 to 20.