Electronic device and method for wireless communication, and computer readable storage medium

CN120052016APending Publication Date: 2025-05-27SONY GROUP CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380072426.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-25
Filing Date
2023-11-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In wireless communications, during the federated reinforcement learning (FRL) process, privacy issues of user information leakage and data transmission congestion and delay issues are difficult to effectively solve, especially when multiple users participate in model updates, which affects the efficiency of model aggregation and distribution. .

Method used

By optimizing the user selection and model upload coordination methods for participating in FRL, using the model dynamic update mechanism, model compression algorithm and optimized model aggregation algorithm, high-quality wireless communication terminals are selected to participate in learning, and auxiliary model upload is performed through P2P communication to reduce the cost of data transmission. Resource usage and delay.

Benefits of technology

It effectively reduces the probability of data transmission congestion, reduces transmission delay, and at the same time ensures learning performance and timely update and convergence speed of the global model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120052016A_ABST
    Figure CN120052016A_ABST
Patent Text Reader

Abstract

The invention provides an electronic device and method for wireless communication and a computer readable storage medium, and the electronic device comprises a processing circuit which is configured to determine a wireless communication terminal to participate in federal reinforcement learning at least based on the processing capability and the wireless communication environment of the wireless communication terminal in the coverage range of a wireless receiving and transmitting node; and acquiring respective local learning models from the wireless communication terminals participating in federal reinforcement learning, and acquiring an updated global model based on the local learning models.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for wireless communication, and computer-readable storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 25, 2022, with application number 202211488978.8 and invention name “Electronic device and method for wireless communication, computer-readable storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of wireless communication technology, and more particularly to updating a global model based on federated reinforcement learning. More particularly, the present application relates to an electronic device and method for wireless communication and a computer-readable storage medium. Background Art

[0003] As machine learning develops, it becomes more capable of solving more complex problems, such as image processing, language recognition, and semantic understanding.

[0004] Technically, federated learning (FL) is a distributed, joint learning solution that leverages multiple users to train their own local data to jointly build a shared model while maintaining the privacy of user data. Figure 1 shows an example FL model, where each participant trains a local model based on their own dataset, such as local models A to C, and submits these local models to the coordinator, which aggregates them to form a global model. The coordinator also provides the updated global model to each participant.

[0005] Reinforcement Learning (RL) is a branch of machine learning that focuses on how individual users interact with the environment and maximize cumulative rewards. The reinforcement learning process allows individuals to learn and improve their behavior through error-tolerant trial and error. Through a series of strategies, individuals participating in reinforcement learning take actions to explore the environment and expect to receive corresponding rewards. Figure 2 shows an example of a reinforcement learning model, where the agent acting as an individual user is in state S t Take action based on the environment t , and in the next state S t+1 Get reward R t+1 For reinforcement learning, a very important issue is to avoid the leakage of user information in order to protect user privacy to the greatest extent possible, because the transmission of raw data between individuals and the central processor will expose great security risks.

[0006] The advantages of federated learning are evident at this point. Not only does it enable information exchange without leaking user privacy, but it also helps users adapt to different environments. Another challenge with reinforcement learning is that many algorithms require pre-training models in a simulated environment, which cannot fully reflect or replicate the real world. Federated learning, however, can bridge the gap between simulation and real-world environments.

[0007] Based on this, the concept of federated reinforcement learning (FRL) came into being. In other words, federated reinforcement learning can be seen as a combination of federated learning and reinforcement learning under data privacy protection. Some reinforcement learning parameters can be presented in federated learning and handle continuous decision-making tasks.

[0008] Summary of the Invention

[0009] A brief overview of the present disclosure is provided below to provide a basic understanding of certain aspects of the present disclosure. It should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify key or important aspects of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is simply to present certain concepts in a simplified form as a prelude to the more detailed description discussed later.

[0010] According to one aspect of the present disclosure, an electronic device for wireless communication is provided, including a processing circuit, which is configured to: determine the wireless communication terminals to participate in federated reinforcement learning based at least on the processing capabilities and wireless communication environment of the wireless communication terminals within the coverage area of ​​the wireless transceiver node; and obtain respective local learning models from the wireless communication terminals participating in the federated reinforcement learning, and obtain an updated global model based on the local learning models.

[0011] According to another aspect of the present disclosure, a method for wireless communication is provided, including: determining the wireless communication terminals to participate in federated reinforcement learning based at least on the processing capabilities and wireless communication environment of the wireless communication terminals within the coverage area of ​​the wireless transceiver node; and obtaining respective local learning models from the wireless communication terminals participating in the federated reinforcement learning, and obtaining an updated global model based on the local learning models.

[0012] According to one aspect of the present disclosure, an electronic device for wireless communication is provided, including a processing circuit configured to: perform training of a local learning model at a wireless communication terminal in response to confirmation information from a wireless transceiver node, wherein the confirmation information indicates that the wireless transceiver node determines that the wireless communication terminal is to participate in federated reinforcement learning based on the processing capability of the wireless communication terminal and the wireless communication environment; and upload the local learning model to the wireless transceiver node and obtain an updated global model from the wireless transceiver node.

[0013] According to another aspect of the present disclosure, a method for wireless communication is provided, comprising: performing training of a local learning model at a wireless communication terminal in response to confirmation information from a wireless transceiver node, wherein the confirmation information indicates that the wireless transceiver node determines that the wireless communication terminal is to participate in federated reinforcement learning based on the processing capability of the wireless communication terminal and the wireless communication environment; and uploading the local learning model to the wireless transceiver node and obtaining an updated global model from the wireless transceiver node.

[0014] According to other aspects of the present disclosure, a computer program code and a computer program product for implementing the above-mentioned method for wireless communication, as well as a computer-readable storage medium having the computer program code for implementing the above-mentioned method for wireless communication recorded thereon, are also provided.

[0015] The electronic device and method according to the embodiments of the present application can selectively determine the wireless communication terminals to participate in federated reinforcement learning, reduce the probability of data transmission congestion, reduce transmission delay, and ensure learning performance.

[0016] These and other advantages of the present disclosure will become more apparent through the following detailed description of the preferred embodiments of the present disclosure in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to further illustrate the above and other advantages and features of the present disclosure, the following is a further detailed description of the specific embodiments of the present disclosure in conjunction with the accompanying drawings. The drawings, together with the detailed description below, are included in this specification and form a part of this specification. Elements with the same function and structure are represented by the same reference numerals. It should be understood that these drawings only depict typical examples of the present disclosure and should not be regarded as limiting the scope of the present disclosure. In the drawings:

[0018] Figure 1 shows an example model of federated learning;

[0019] Figure 2 shows an example of a reinforcement learning model;

[0020] FIG3 is a block diagram showing functional modules of an electronic device for wireless communication according to an embodiment of the present application;

[0021] FIG4 shows an example of the flow of information related to the FRL process;

[0022] FIG5 shows another example of the flow of related information for the FRL process;

[0023] Figure 6 shows an example of the relevant information flow of auxiliary model upload in FRL;

[0024] FIG7 is a block diagram showing functional modules of an electronic device for wireless communication according to another embodiment of the present application;

[0025] FIG8 shows another example of the related information flow of auxiliary model upload in FRL;

[0026] Figure 9 shows a schematic diagram of the reward function design;

[0027] FIG10 shows an example of the related information flow of FRL in a fleet application scenario;

[0028] FIG11 shows a flowchart of a method for wireless communication according to an embodiment of the present application;

[0029] FIG12 shows a flowchart of a method for wireless communication according to another embodiment of the present application;

[0030] FIG13 is a block diagram showing a first example of a schematic configuration of an eNB or gNB to which the technology of the present disclosure may be applied;

[0031] FIG14 is a block diagram illustrating a second example of a schematic configuration of an eNB or gNB to which the technology of the present disclosure may be applied;

[0032] FIG15 is a block diagram showing an example of a schematic configuration of a smartphone to which the technology of the present disclosure can be applied;

[0033] FIG16 is a block diagram showing an example of a schematic configuration of a car navigation device to which the technology of the present disclosure can be applied; and

[0034] FIG17 is a block diagram of an exemplary structure of a general-purpose personal computer in which the method and / or apparatus and / or system according to the embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0035] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings. For the sake of clarity and conciseness, not all features of an actual implementation are described in this specification. However, it should be understood that in the process of developing any such actual implementation, many implementation-specific decisions must be made in order to achieve the developer's specific goals, such as compliance with system and business-related constraints, which may vary from implementation to implementation. In addition, it should be understood that although the development work may be very complex and time-consuming, it is a routine task for those skilled in the art who benefit from the contents of this disclosure.

[0036] It is also necessary to explain here that, in order to avoid obscuring the present disclosure due to unnecessary details, the accompanying drawings only show the device structure and / or processing steps that are closely related to the solution according to the present disclosure, while other details that are not closely related to the present disclosure are omitted.

[0037] <First embodiment>

[0038] As mentioned above, FRL can be used to train global models in a network environment. In this case, users using FRL need to interact with other users and / or network-side servers for a large number of model parameters, intermediate results, etc., which will consume a lot of communication resources and electricity. Therefore, the overhead of communication resources and the battery capacity limitations of user-side devices are issues that must be considered. For example, it is necessary to coordinate the corresponding model upload and download strategies to maximize the efficiency of model update iterations. Specifically, for example, a model dynamic update mechanism can be used to optimize the number of model interactions, a model compression algorithm can be used to reduce the size of interaction data, and the model aggregation algorithm can be optimized to allow users participating in FRL to upload only important parameters of the local learning model, etc.

[0039] In this embodiment, a technical solution will be provided to improve the efficiency of model update iteration by optimizing the selection of users participating in FRL, or also optimizing the model upload coordination method between users.

[0040] Figure 3 shows a functional module block diagram of an electronic device 100 for wireless communication according to this embodiment. As shown in Figure 3, the electronic device 100 includes: a determination unit 101, configured to determine the wireless communication terminals to participate in FRL based at least on the processing capabilities and wireless communication environment of the wireless communication terminals within the coverage range of the wireless transceiver node; a communication unit 102, configured to obtain respective local learning models from the wireless communication terminals participating in FRL; and an acquisition unit 103, configured to obtain an updated global model based on the local learning model.

[0041] The electronic device 100 can be provided on the network side or the cloud server side, for example, on the wireless transceiver node side. The wireless transceiver node here can be a base station, an access point (AP), a roadside unit, etc. In addition, the wireless communication terminal here can be various user equipment (UE) or user terminals capable of participating in FRL, or a communication terminal such as a mobile base station.

[0042] The determining unit 101, the communicating unit 102, and the acquiring unit 103 may be implemented by one or more processing circuits, which may be implemented as chips or processors, for example. Furthermore, it should be understood that the various functional units in the electronic device shown in FIG3 are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementations.

[0043] It should also be noted that the electronic device 100 can be implemented at the chip level or at the device level. For example, the electronic device 100 can operate as a wireless transceiver node itself, and may also include external devices such as a memory and a transceiver (not shown in the figure). The memory can be used to store programs and related data information that need to be executed by the wireless transceiver node to implement various functions. The transceiver may include one or more communication interfaces to support communication with different devices (e.g., other wireless transceiver nodes, UE, etc.), and the implementation form of the transceiver is not specifically limited here.

[0044] In theory, the more wireless communication terminals participating in FRL, the better the learning performance. This is because a richer dataset enhances the system model's understanding of the learning object. However, after local learning, wireless communication terminals need to upload model parameters, which consumes wireless resources. The more wireless communication terminals participating in FRL, the more uplink transmission resources are simultaneously occupied, which can easily lead to network storms for data uploads, causing data congestion and increased transmission delays, which in turn seriously affects the effectiveness of model aggregation and the timeliness of model distribution.

[0045] In this embodiment, the determining unit 101 selects wireless communication terminals to participate in FRL to minimize the number of wireless communication terminals participating in FRL while ensuring learning performance. Specifically, the determining unit 101 may consider the processing capabilities of the wireless communication terminals and the current wireless communication environment to select the most suitable wireless communication terminal to participate in FRL.

[0046] For example, the wireless communication environment of a wireless communication terminal may include one or more of the following: wireless channel quality, data rate, interference intensity, geographic location, information transmission path loss related to the geographic location, and mobility speed. Determination unit 101 may preferentially determine that wireless communication terminals with better wireless communication environments participate in FRL to ensure transmission of model parameters. Similarly, determination unit 101 may preferentially determine that wireless communication terminals with stronger processing capabilities participate in FRL to ensure effective local model learning and model parameter upload.

[0047] For example, when determining learning terminals for FRL, the wireless communication environment of the wireless communication terminal is weighted more highly than the processing capability of the wireless communication terminal. In other words, when determining unit 101 cannot strike a balance between selecting a wireless communication terminal with the best possible wireless communication environment and a wireless communication terminal with the highest possible processing capability, the wireless communication terminal with the best possible wireless communication environment is prioritized to ensure that the model update latency meets the requirements. It should be understood that this is not restrictive and can be modified based on actual needs.

[0048] The communication unit 102 may at least partially obtain information about the wireless communication environment and processing capabilities from the wireless communication terminal. For example, the network where the electronic device 100 resides may obtain its own mobile speed and obtain its geographic location, channel quality, processing capabilities, etc. through reports from the wireless communication terminal. When the determination unit 101 is able to determine whether to select certain wireless communication terminals for FRL based on the network's own information, it may disable these wireless communication terminals from reporting channel status in a certain round to reduce signaling overhead.

[0049] In an exemplary scenario based on the Internet of Vehicles (IoV), the wireless communication terminal can be a vehicle or a vehicle in a platoon, and the learning object can be, for example, the vehicle's environment, such as surrounding traffic conditions. The processing capabilities of the wireless communication terminal can include, for example, one or more of the following: the vehicle's automation level and its environmental awareness capabilities. It should be noted that in some of the following descriptions, IoV will be used as an example scenario; however, it should be understood that this is not restrictive.

[0050] The wireless communication terminals participating in FRL train local learning models based on their own stored data sets, and send the trained local learning models to the electronic device 100 for model aggregation. Note that the wireless communication terminals participating in FRL initially receive the distributed initial machine learning model from the network side and perform initial training based on the initial machine learning model. The acquisition unit 103, for example, aggregates the local learning models obtained from each wireless communication terminal to obtain an updated global model. The acquisition unit 103 can, for example, perform the function of the coordinator shown in Figure 1.

[0051] In the connected vehicle scenario, vehicles can set a predetermined reward function during local learning model training. This reward function may include, for example, one or more of the following: a forward driving reward, a collision avoidance reward, and a speed maintenance reward. These reward functions are described in detail below.

[0052] In one example, the communication unit 102 is further configured to provide an updated global model to wireless communication terminals within the coverage of the wireless transceiver node. This enables all wireless communication terminals to be informed of the latest global model update. Because before the model converges, the number and identification of wireless communication terminals participating in FRL determined in each round of learning may be different due to factors such as changes in the wireless communication environment. In this example, even if a wireless communication terminal has not been selected to participate in FRL for a long time, it can obtain the latest updated global model, so that if it is suddenly selected to participate in FRL, it can also be trained based on the latest global model, so that it can contribute positive feedback to the global model when the model is aggregated, which is conducive to improving the convergence speed of the global model.

[0053] For example, the communication unit 102 may provide a complete updated global model to wireless communication terminals participating in FRL, and provide a lightweight updated global model to wireless communication terminals not participating in FRL, wherein the lightweight global model is, for example, a global model with low complexity or low precision.

[0054] To facilitate understanding, FIG4 shows an example of the flow of related information of the FRL process, wherein the wireless communication terminal is referred to as the terminal, and the network side performs the functions of the above-mentioned electronic device 100. It should be noted that in the information flow shown in FIG4 and subsequent figures, the number of terminals is only exemplary and is not limited to the three shown in the figure.

[0055] First, the network distributes the initial machine learning model to each terminal A through C as the basis for their training. Next, the first round of FRL begins, with potential FRL participants A through C reporting their own information to the network, such as information about the wireless communication environment and / or processing capabilities. Based on the received terminal-related information and / or its own stored information, the network determines which terminals will participate in FRL. In the example of Figure 4 , terminals A and C are determined to participate in this round of FRL, and the network notifies terminals A and C of the confirmation information. After receiving the confirmation information, terminals A and C train local learning models based on their local datasets and the initial machine learning model, and upload the local learning models to the network. The network aggregates the received local learning models to obtain an updated global model, and sends the complete updated global model to terminals A and C participating in FRL, and sends a lightweight updated global model to terminal B not participating in FRL. This ensures fairness while ensuring that each terminal maintains the latest version of the global model.

[0056] If the model has not converged, the second round of FRL continues. Similarly, potential FRL participants A through C each report relevant information to the network, such as information about the wireless communication environment and / or processing capabilities. Note that if the network determines, for example, that Terminal C does not report relevant information, it can still determine whether Terminal C should participate in this round of FRL and disable Terminal C from reporting information in this round of FRL to reduce signaling overhead. The network determines the terminals to participate in FRL based on received terminal-related information and / or its own stored information. In the example of Figure 4, Terminals A and B are determined to participate in this round of FRL, and the network notifies Terminals A and B of the confirmation information. After receiving the confirmation information, Terminals A and B train local learning models based on their local datasets and the global model obtained in the first round, and upload the local learning models to the network. The network aggregates the received local learning models to obtain an updated global model, and sends the complete updated global model to Terminals A and B participating in FRL, and sends a lightweight updated global model to Terminal C not participating in FRL. For example, the network side determines whether the updated global model has converged. If not, the next round of FRL is continued; otherwise, the process ends.

[0057] In addition, the communication unit 102 may also provide an updated global model to wireless communication terminals participating in FRL, wherein wireless communication terminals not participating in FRL obtain the updated global model from wireless communication terminals participating in FRL through P2P communication. An example of the relevant information flow in this case is shown in FIG5 . It can be seen that the only difference between FIG5 and FIG4 is the different way in which the global model of the terminals not participating in FRL is updated. For the sake of brevity, the description of the parts that are the same as FIG4 will not be repeated here. After the first round of FRL, the network side provides the complete updated global model to terminals A and C participating in FRL. Terminal B not participating in FRL sends a global model update request to terminal A through P2P communication. In response to the request, terminal A provides the updated global model obtained from the network side to terminal B. Here, P2P communication can be implemented in a variety of ways, including but not limited to WiFi, Bluetooth, RFID, direct communication link (sidelink) communication, etc. Similarly, after the second round of FRL, terminal C not participating in FRL obtains the updated global model from terminal B through P2P communication.

[0058] The communication unit 102 may also be configured to instruct the wireless communication terminals participating in FRL to perform model distillation to control the data volume of the local learning model to a predetermined amount. In this way, the amount of data of the local learning model to be uploaded can be effectively reduced, thereby reducing latency and transmission resource requirements. This instruction can be sent together with the confirmation information shown in FIG4, or can be sent as a separate signaling, which is not restrictive.

[0059] The predetermined amount may be the same or different for different wireless communication terminals. For example, if the data volume requirements for the distilled models are the same (i.e., the predetermined amount is the same), the predetermined amount may be set to the amount of model data that can be transmitted by the wireless communication terminal with the worst wireless communication environment (e.g., channel quality).

[0060] On the other hand, when the data volume requirements of the models after distillation are different (i.e., the predetermined amounts are different), for example, a model with a large data volume can be configured for a wireless communication terminal with strong processing capabilities, while a model with a small data volume can be configured for a wireless communication terminal with weak processing capabilities. In this case, during the training process of the local learning model, wireless communication terminals with different processing capabilities can exchange their respective underlying training results through P2P communication to ensure that models of different sizes can be aggregated on the network side. P2P communication here includes, but is not limited to, Sidelink communication, WiFi, Bluetooth, RFID, etc.

[0061] For example, the communication unit 102 is further configured to schedule wireless communication terminals participating in FRL to upload local learning models so that the end time of model upload by each wireless communication terminal is consistent. For example, the communication unit 102 can schedule based on the data size of the model configured for the wireless communication terminal.

[0062] In summary, the electronic device 100 according to this embodiment can selectively determine the wireless communication terminals to participate in federated reinforcement learning, thereby reducing the probability of data transmission congestion and transmission latency while ensuring learning performance. In addition, in this embodiment, all wireless communication terminals within the coverage area of ​​the wireless transceiver node can maintain the update of the global model, thereby ensuring that positive feedback to the global model can be provided when participating in FRL, which is conducive to improving the convergence speed of the global model.

[0063] <Second embodiment>

[0064] Since the wireless communication environment may change at any time, it is possible that wireless communication terminals participating in FRL may not be able to upload the local learning model in a timely manner due to sudden changes in the wireless communication environment. In order to ensure timely updates of the global model, this embodiment adopts a solution in which other wireless communication terminals participating in FRL perform auxiliary communication.

[0065] The electronic device 100 according to this embodiment includes the same functional modules as the electronic device 100 in the first embodiment. For example, the communication unit 102 is also configured to schedule the first wireless communication terminal participating in the FRL to assist the second wireless communication terminal participating in the FRL in uploading the local learning model, wherein the second wireless communication terminal sends part or all of the local learning model to the first wireless communication terminal through P2P communication for uploading by the first wireless communication terminal. The P2P communication here includes, for example, but is not limited to Sidelink communication, WiFi, Bluetooth, RFID, etc. Among them, the first and the second are only for the purpose of distinction, and do not represent any meaning of any order or priority.

[0066] For example, the wireless communication environment of the first wireless communication terminal is better than that of the second wireless communication terminal. In this case, the determination unit 101 determines, for example, that the channel state of the first wireless communication terminal is suitable for transmitting a large model, while the channel state of the second wireless communication terminal is only suitable for transmitting a small model or is unsuitable for uploading a model. Thus, the first wireless communication terminal can assist the second wireless communication terminal in uploading the model.

[0067] For ease of understanding, Figure 6 shows an example of the relevant information flow for assisting model upload in FRL. Figure 6 shows the information flow of a round of FRL, in which the network side determines that terminal A and terminal C are to participate in this round of FRL. In the example of Figure 6, it is assumed that the wireless communication environment (such as the channel state) of terminal A is better than that of terminal C, so the network side determines that terminal A will assist terminal C in uploading the model. Therefore, the network side sends corresponding scheduling information while sending confirmation information to terminal A and terminal C. The scheduling information, for example, instructs terminal A to assist terminal C in uploading the local learning model. It should be noted that although the confirmation information and scheduling information are shown as one signaling in Figure 6, this is not restrictive. The two can be sent separately, and the timing of sending the scheduling information is not limited to this. For example, it can be sent during the training or model upload process of the local learning model. In addition, the scheduling information can also be sent in response to the request of terminal C, which is not restrictive.

[0068] The scheduling information may include, for example, one or more of the following: the identifiers (IDs) of terminal A and terminal C, the size of the model to be uploaded, an incentive strategy, a P2P communication method, etc. Among them, the incentive strategy may be, for example, additional rewards provided to encourage terminals to assist other terminals in uploading models, such as additional model sharing, etc. As mentioned above, the P2P communication here includes, for example, but is not limited to, Sidelink communication, WiFi, Bluetooth, RFID, etc. In addition, the scheduling information for terminal A providing assistance and terminal C receiving assistance may have different contents. For example, the scheduling information for terminal A includes an instruction to provide assistance to terminal C, while the scheduling information for terminal C includes an instruction to indicate that terminal A will provide assistance.

[0069] Terminals A and C train their local learning models separately. After training is complete, terminal C sends (shares) all or part of the trained local learning model to terminal A via P2P communication based on scheduling information. Terminal A then uploads the model along with its own local learning model to the network. Because terminal A has a good channel condition, it can achieve faster model uploads, reducing latency. If terminal C only shares a portion of the local learning model, it uploads the remaining portion to the network.

[0070] Furthermore, when the P2P communication is sidelink communication, terminal C can establish a sidelink between terminal C and terminal A in response to the aforementioned scheduling information. As shown by the dashed line in Figure 6 , terminal C sends a sidelink establishment request to terminal A through the direct terminal discovery process. Terminal A responds to this request by sending a sidelink establishment confirmation to terminal C, thereby establishing a sidelink between the two. A portion or all of terminal C's locally learned model is transmitted to terminal A via the sidelink.

[0071] To sum up, according to this embodiment, the electronic device 100 schedules the wireless communication terminals participating in FRL to upload the auxiliary model based on P2P communication, which further reduces the delay caused by model uploading, ensures the timely update of the global model, and is conducive to improving the convergence speed of the global model.

[0072] <Third embodiment>

[0073] Figure 7 shows a functional module block diagram of an electronic device 200 for wireless communication according to another embodiment of the present application. As shown in Figure 7, the electronic device 200 includes: a training unit 201, configured to perform training of a local learning model at a wireless communication terminal in response to confirmation information from a wireless transceiver node, wherein the confirmation information indicates that the wireless transceiver node determines that the wireless communication terminal is to participate in FRL based on the processing capability of the wireless communication terminal and the wireless communication environment; and a communication unit 202, configured to upload the local learning model to the wireless transceiver node and obtain an updated global model from the wireless transceiver node.

[0074] The training unit 201 and the communication unit 202 may be implemented by one or more processing circuits, such as chips or processors. Furthermore, it should be understood that the various functional units in the electronic device shown in FIG7 are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementations.

[0075] The electronic device 200 is, for example, disposed on a wireless communication terminal side or communicatively connected to a wireless communication terminal. The wireless communication terminal here can be various user equipment (UE) or user terminals capable of participating in FRL, or a communication terminal such as a mobile base station. The wireless transceiver node here can be a base station, access point (AP), roadside unit, etc., and more generally represents the network side or cloud server side.

[0076] The electronic device 200 can be implemented at the chip level or at the device level. For example, the electronic device 200 can operate as a wireless communication terminal itself and may also include external devices such as a memory and a transceiver (not shown). The memory can be used to store programs and related data information that the wireless communication terminal needs to execute to implement various functions. The transceiver may include one or more communication interfaces to support communication with different devices (e.g., other wireless communication terminals, base stations, core networks, etc.). The implementation form of the transceiver is not specifically limited here.

[0077] Similar to the first embodiment, the wireless communication environment of the wireless communication terminal includes one or more of the following: wireless channel quality, data rate, interference intensity, geographical location, information transmission path loss related to the geographical location, and moving speed.

[0078] In an example scenario of the Internet of Vehicles, the wireless communication terminal is a vehicle or a vehicle in a fleet, and the processing capability of the vehicle includes, for example, one or more of the following: the vehicle's autonomous driving level, and the vehicle's environmental perception capability.

[0079] The communication unit 202 may also be configured to provide the wireless transceiver node with at least a portion of the information about the processing capabilities of the wireless communication terminal and the wireless communication environment. The wireless transceiver node determines, based on this information and / or related information stored in the node, that the wireless communication terminal is to participate in the FRL and sends corresponding confirmation information. The communication unit 202 is configured to receive the confirmation information.

[0080] As a wireless communication terminal participating in FRL, its communication unit 202 can obtain an updated global model from a wireless transceiver node. In one example, the communication unit 202 is also configured to provide an updated global model to a wireless communication terminal that does not participate in FRL through P2P communication. For example, the communication unit 202 receives a global model update request sent through P2P communication from a wireless communication terminal that does not participate in FRL, and in response to the request, provides the obtained updated global model to the wireless communication terminal through P2P communication. Here, P2P communication can be implemented in a variety of ways, including but not limited to WiFi, Bluetooth, RFID, sidelink communication, etc. An example of a detailed information flow has been given in Figure 5 and will not be repeated here.

[0081] In addition, for example, when the wireless communication environment of the wireless communication terminal is poor, the communication unit 202 can also be configured to upload part or all of the local learning model with the assistance of other wireless communication terminals participating in FRL.

[0082] As an example, the communication unit 202 may receive scheduling information from a wireless transceiver node, and based on the scheduling information, perform model transmission between the wireless communication terminal and the other wireless communication terminal through P2P communication. Here, P2P communication can be implemented in a variety of ways, including but not limited to WiFi, Bluetooth, RFID, sidelink communication, and the like. For example, the communication unit 202 may establish a sidelink between the wireless communication terminal and the other wireless communication terminal in response to the scheduling information. The scheduling information may include, for example, one or more of the following: identifications (IDs) of the wireless communication terminal and the other wireless communication terminal, the size of the model to be uploaded, incentive strategies, P2P communication methods, and the like. The relevant details are given in detail with reference to FIG6 in the second embodiment, and are also used here and will not be repeated.

[0083] As another example, the communication unit 202 can determine other wireless communication terminals that can provide assistance by broadcasting request information, and the wireless communication terminal performs P2P communication with the other wireless communication terminals. The request information includes, for example, one or more of the following: the ID of the wireless communication terminal, the size of the model to be uploaded, and the incentive strategy. For example, after receiving the confirmation information, the communication unit 202 begins to look for other wireless communication terminals in the surrounding area that are willing and capable of assisting it in uploading the model to further reduce latency. After finding such other wireless communication terminals, after the training of the local learning model is completed, the communication unit 202 sends part or all of the local learning model to the other wireless communication terminal through P2P communication.

[0084] For ease of understanding, Figure 8 illustrates the information flow for assisting model upload in FRL according to this example. Figure 8 illustrates the information flow for a round of FRL. The process prior to determining the terminals participating in the FRL is identical to the corresponding portion of Figure 6 and will not be repeated here. In Figure 8, the network sends confirmation information to terminals A and C, which have been determined to participate in the FRL. Assume that terminal C, which receives the confirmation information, experiences a poor wireless communication environment and therefore requires P2P communication to upload the assisting model. Therefore, terminal C broadcasts a request for assisting model upload within its coverage area. This request may include, for example, one or more of the following: terminal C's ID, the size of the model to be uploaded, and an incentive policy. After receiving this broadcast request, terminal A determines to provide assistance to terminal C and sends a confirmation of the assistance upload request. Terminal C trains its local learning model and, upon completion, transmits (shares) a portion or all of the local learning model to terminal A via P2P communication. Terminal A then uploads this portion of the model along with its own local learning model to the network. If terminal C only shares a portion of the local learning model, terminal C itself uploads the remaining portion to the network. Similarly, P2P communication can be implemented in a variety of ways, including but not limited to WiFi, Bluetooth, RFID, sidelink communication, etc.

[0085] According to this embodiment, the electronic device 200 utilizes other wireless communication terminals participating in FRL to perform auxiliary model upload based on P2P communication, thereby reducing the delay caused by model upload in a poor wireless communication environment, ensuring timely updating of the global model, and helping to improve the convergence speed of the global model.

[0086] In an example scenario of the Internet of Vehicles (IoV), the wireless communication terminal is a vehicle or a vehicle in a fleet. The training unit 201 is configured to set a predetermined reward function when training the local learning model. The reward function may include, for example, one or more of the following: a forward driving reward, a collision avoidance reward, and a speed maintenance reward.

[0087] The main purpose of reward function design is to judge the effectiveness of the vehicle's behavior based on its behavior and response to the current environment, and ultimately enable the vehicle to continuously learn and iterate its behavior towards achieving its goals, as shown in the schematic diagram in Figure 9, where S0 to S3 represent states, and the numbers above the arrows represent reward values ​​or penalty values. Positive numbers represent reward values, and negative numbers represent penalty values.

[0088] For example, the forward driving reward is mainly used to ensure that the vehicle can maintain a relatively correct direction of travel, including the vehicle's forward direction towards the road and the vehicle does not exceed the lane lines on the left and right sides. For example, the forward driving reward can be calculated as follows:

[0089] Where e1 represents the vehicle's offset distance, e2 represents the vehicle's offset angle, and e′ represents the rate of change of the two parameters. k1k2k3k4 are the corresponding variable coefficients.

[0090] The collision avoidance reward is mainly used to judge the safe distance between the vehicle and the obstacle, so as to avoid the vehicle colliding with it. Refer to the following formula (2), which is based on the Berkeley Algorithm to calculate the dynamic safety distance (LoSD(v f ), LaSD(v avut ), and the vehicle's speed v f and v avut , acceleration α f and α avut , the vehicle steering angle β, and the vehicle response time τ, where the subscript avut represents the current vehicle in the test and the subscript f represents the following vehicle (or preceding vehicle). By calculating the safe distance and obtaining location information from the server, the probability of a collision can be calculated. The greater the collision probability, the greater the penalty.

[0091] Among them, R min It represents the minimum distance that the current vehicle needs to maintain when the following vehicle stops suddenly, and also represents the necessary tolerance for distance calculation errors. When the speed is less than or equal to the speed limit, a reward will be generated along the road direction for each distance traveled; if the speed exceeds the speed limit, a penalty will be generated. The following formula (3) shows that after the vehicle travels a valid distance, the reward increases as the speed along the road increases, thereby encouraging the vehicle to accelerate; but if the vehicle speed exceeds the speed limit, a penalty will be generated, causing the vehicle to drive close to the speed limit. Among them, Rv keep represents the speed maintenance reward value, ω represents the reward coefficient for each distance, V t *cosβ represents the speed along the road, v max Indicates the maximum speed of the road.

[0092] It should be understood that the above reward functions are only examples and are not limiting.

[0093] Furthermore, when the reward function is applied to a convoy, because the convoy is considered as a whole during its journey, it is not necessary for every vehicle in the convoy to participate in FRL from a resource consumption perspective. The reward function also applies to vehicles participating in FRL and other member vehicles that are only performing local learning. For example, the pilot vehicle can participate in FRL. Vehicles not participating in FRL can obtain the latest global model from the pilot vehicle to help improve their perception and decision-making capabilities in the corresponding environment.

[0094] Team members who do not participate in FRL can use one of the following two methods to update their driving strategies: first, instead of actively learning the surrounding environment and updating the driving strategy, they can directly obtain vehicle control instructions and corresponding vehicle driving parameters from the pilot car; second, they can actively learn the surrounding environment and update the driving strategy. For example, they can obtain an updated global model from the pilot car and update the vehicle's driving strategy based on their own local learning results.

[0095] In addition, fleet members participating in FRL can send their local learning models to the pilot vehicle. The pilot vehicle can then perform a rough update of the fleet's local model based on these local learning models and send the roughly updated fleet local model to the network. Accordingly, if the wireless communication terminal corresponding to the electronic device 200 is the pilot vehicle in the fleet, the training unit 201 is further configured to perform a rough update of the fleet's local model based on the fleet members' local learning models and upload the roughly updated fleet local model to the wireless transceiver node. Furthermore, the communication unit 202 is further configured to transmit the updated global model obtained from the wireless transceiver node to each fleet member.

[0096] For ease of understanding, Figure 10 shows an example of the relevant information flow for FRL in a fleet application scenario. The information flow before determining the vehicles participating in the FRL is similar to the corresponding process shown in Figure 4 and will not be repeated here. The network side determines that pilot vehicle A and vehicle C will participate in the FRL. However, the network side will send confirmation information for both to pilot vehicle A, and pilot vehicle A will forward the confirmation information (for example, the ID of the vehicle participating in the FRL) to vehicle C. Pilot vehicle A and vehicle C train local learning models based on local datasets. After completion, vehicle C transmits its local learning model to pilot vehicle A through the fleet. Pilot vehicle A performs a rough update of the fleet local model based on its own local learning model and vehicle C's local learning model, and uploads the rough updated fleet local model to the network side. In addition, vehicle C does not need to upload its local learning model, thereby reducing resource consumption. The network side aggregates the uploaded fleet local model and the local learning models obtained from other vehicles to obtain an updated global model and sends it to pilot vehicle A. The pilot car A distributes the updated global model within the fleet so that the fleet members can update the global model synchronously.

[0097] <Fourth embodiment>

[0098] In the process of describing the electronic device for wireless communication in the above embodiments, it is obvious that some processes or methods are also disclosed. Below, an overview of these methods is given without repeating some of the details discussed above, but it should be noted that although these methods are disclosed in the process of describing the electronic device for wireless communication, these methods do not necessarily use the components described or are not necessarily performed by those components. For example, the embodiments of the electronic device for wireless communication can be partially or completely implemented using hardware and / or firmware, and the methods for wireless communication discussed below can be completely implemented by computer-executable programs, although these methods can also use the hardware and / or firmware of the electronic device for wireless communication.

[0099] FIG11 shows a flowchart of a method for wireless communication according to an embodiment of the present application, the method comprising: determining wireless communication terminals to participate in FRL based at least on the processing capabilities and wireless communication environment of wireless communication terminals within the coverage area of ​​a wireless transceiver node (S11); and obtaining local learning models from the wireless communication terminals participating in FRL, and obtaining an updated global model based on the local learning models (S12). The method can be performed, for example, on the wireless transceiver node side.

[0100] For example, the wireless communication environment of a wireless communication terminal includes one or more of the following: wireless channel quality, data rate, interference intensity, geographic location, information transmission path loss related to the geographic location, and mobile speed. Information about the wireless communication environment and processing capabilities can be obtained, at least in part, from the wireless communication terminal. For example, when determining a wireless communication terminal to participate in FRL, the wireless communication environment of the wireless communication terminal can be given a higher weight than the processing capabilities of the wireless communication terminal.

[0101] In addition, as shown in the dotted box in FIG11 , the method further includes step S13: providing an updated global model to wireless communication terminals within the coverage area. As an example, a complete updated global model may be provided to wireless communication terminals participating in the FRL, and a lightweight updated global model may be provided to wireless communication terminals not participating in the FRL. As another example, an updated global model may be provided to wireless communication terminals participating in the FRL, wherein the wireless communication terminals not participating in the FRL obtain the updated global model from the wireless communication terminals participating in the FRL through P2P communication.

[0102] In addition, although not shown in the figure, the above method may also include: scheduling a first wireless communication terminal participating in the FRL to assist a second wireless communication terminal participating in the FRL in uploading a local learning model, wherein the second wireless communication terminal transmits a portion or all of the local learning model to the first wireless communication terminal via P2P communication for uploading by the first wireless communication terminal. For example, the wireless communication environment of the first wireless communication terminal is better than that of the second wireless communication terminal. Exemplarily, the P2P communication includes sidelink communication.

[0103] The above method may also include: instructing the wireless communication terminals participating in FRL to perform model distillation to control the data volume of the local learning model to a predetermined amount. The above method may also include: scheduling the wireless communication terminals participating in FRL to upload the local learning model so that the end time of the model upload of each wireless communication terminal remains consistent.

[0104] As an example, the wireless communication terminal can be a vehicle or a vehicle in a fleet. The vehicle's processing capabilities include, for example, one or more of the following: the vehicle's autonomous driving level and the vehicle's environmental perception capabilities. When training a local learning model, the vehicle can set a predetermined reward function, which includes, for example, one or more of the following: a forward driving reward, a collision avoidance reward, and a speed maintenance reward.

[0105] The above method corresponds to the electronic device 100 in the first and second embodiments. The relevant detailed description has been given in the first and second embodiments and will not be repeated here.

[0106] FIG12 shows a flowchart of a method for wireless communication according to another embodiment of the present application, the method comprising: in response to confirmation information from a wireless transceiver node, performing training of a local learning model at a wireless communication terminal (S21), wherein the confirmation information indicates that the wireless transceiver node determines that the wireless communication terminal is to participate in FRL based on the processing capability of the wireless communication terminal and the wireless communication environment; and uploading the local learning model to the wireless transceiver node and obtaining an updated global model from the wireless transceiver node (S22). This method can be performed, for example, on the wireless communication terminal side.

[0107] Similarly, the wireless communication environment of the wireless communication terminal may include one or more of the following: wireless channel quality, data rate, interference intensity, geographical location, information transmission path loss related to the geographical location, and moving speed.

[0108] Although not shown in the figure, the above method may further include the following step: providing an updated global model to a wireless communication terminal that does not participate in the FRL in response to a request from the terminal through P2P communication.

[0109] The above method also includes: uploading a part or all of the local learning model of the wireless communication terminal with the assistance of other wireless communication terminals participating in FRL.

[0110] In one example, other wireless communication terminals may be identified by broadcasting request information, and P2P communication may be performed between the wireless communication terminal and the other wireless communication terminals. For example, the request information may include one or more of the following: an identifier of the wireless communication terminal, the size of the model to be uploaded, and an incentive strategy.

[0111] In another example, the model may be transmitted between a wireless communication terminal and other wireless communication terminals through P2P communication based on scheduling information from a wireless transceiver node.

[0112] For example, the wireless communication terminal may be a vehicle or a vehicle in a fleet. The processing capability of the vehicle includes one or more of the following: the autonomous driving level of the vehicle, and the environmental perception capability of the vehicle.

[0113] The above method also includes: when the vehicle is training the local learning model, setting a predetermined reward function, the reward function including one or more of the following: forward driving reward, collision avoidance reward and speed maintenance reward.

[0114] If the wireless communication terminal is a pilot vehicle in a fleet, the method further includes: the pilot vehicle coarsely updating the fleet's local learning model based on the local learning models of the fleet members, and uploading the coarsely updated fleet local learning model to the wireless transceiver node. The pilot vehicle may also transmit the updated global model obtained from the wireless transceiver node to each fleet member.

[0115] The above method corresponds to the electronic device 200 in the third embodiment. The relevant detailed description has been given in the third embodiment and will not be repeated here.

[0116] Note that the above methods can be used in combination or individually.

[0117] The technology of the present disclosure can be applied to various products.

[0118] For example, the electronic device 100 can also be implemented as various base stations. The base station can be implemented as any type of evolved Node B (eNB) or gNB (5G base station). eNBs include, for example, macro eNBs and small eNBs. Small eNBs can be eNBs that cover cells smaller than macro cells, such as pico eNBs, micro eNBs, and home (femto) eNBs. Similar situations can also apply to gNBs. Alternatively, the base station can be implemented as any other type of base station, such as a NodeB and a base transceiver station (BTS). The base station may include: a main body (also referred to as a base station device) configured to control wireless communications; and one or more remote radio heads (RRHs) located at a different place from the main body. In addition, various types of user equipment can work as a base station by temporarily or semi-permanently performing base station functions.

[0119] The electronic device 200 can be implemented as various user devices. The user device can be implemented as a mobile terminal (such as a smartphone, a tablet personal computer (PC), a notebook PC, a portable gaming terminal, a portable / dongle-type mobile router, and a digital camera) or an in-vehicle terminal (such as a car navigation device). The user device can also be implemented as a terminal that performs machine-to-machine (M2M) communication (also known as a machine-type communication (MTC) terminal). In addition, the user device can be a wireless communication module (such as an integrated circuit module consisting of a single chip) installed in each of the above terminals.

[0120] [Application examples for base stations]

[0121] (First application example)

[0122] FIG13 is a block diagram illustrating a first example of a schematic configuration of an eNB or gNB to which the techniques of this disclosure can be applied. Note that the following description uses an eNB as an example, but is equally applicable to a gNB. An eNB 800 includes one or more antennas 810 and a base station device 820. The base station device 820 and each antenna 810 can be connected to each other via an RF cable.

[0123] Each of the antennas 810 includes a single or multiple antenna elements (such as multiple antenna elements included in a multiple-input multiple-output (MIMO) antenna) and is used for base station device 820 to transmit and receive wireless signals. As shown in FIG13 , eNB 800 may include multiple antennas 810. For example, multiple antennas 810 may be compatible with multiple frequency bands used by eNB 800. Although FIG13 shows an example in which eNB 800 includes multiple antennas 810, eNB 800 may also include a single antenna 810.

[0124] The base station device 820 includes a controller 821 , a memory 822 , a network interface 823 , and a wireless communication interface 825 .

[0125] The controller 821 may be, for example, a CPU or a DSP, and operates various functions of the higher layers of the base station device 820. For example, the controller 821 generates data packets based on the data in the signal processed by the wireless communication interface 825, and transmits the generated packets via the network interface 823. The controller 821 may bundle data from multiple baseband processors to generate bundled packets, and transmit the generated bundled packets. The controller 821 may have logic functions for performing the following controls: the control may be radio resource control, radio bearer control, mobility management, admission control, and scheduling. The control may be performed in conjunction with a nearby eNB or core network node. The memory 822 includes RAM and ROM, and stores programs executed by the controller 821 and various types of control data (such as a terminal list, transmission power data, and scheduling data).

[0126] The network interface 823 is a communication interface for connecting the base station device 820 to the core network 824. The controller 821 can communicate with the core network node or another eNB via the network interface 823. In this case, the eNB 800 and the core network node or other eNBs can be connected to each other through a logical interface (such as an S1 interface and an X2 interface). The network interface 823 can also be a wired communication interface or a wireless communication interface for a wireless backhaul line. If the network interface 823 is a wireless communication interface, the network interface 823 can use a higher frequency band for wireless communication than the frequency band used by the wireless communication interface 825.

[0127] The wireless communication interface 825 supports any cellular communication scheme, such as Long Term Evolution (LTE) and LTE-Advanced, and provides wireless connectivity to terminals located in the cell of the eNB 800 via the antenna 810. The wireless communication interface 825 may typically include, for example, a baseband (BB) processor 826 and RF circuitry 827. The BB processor 826 can perform various signal processing functions, such as encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and performs various types of signal processing for layers such as Layer 1 (L1), Medium Access Control (MAC), Radio Link Control (RLC), and Packet Data Convergence Protocol (PDCP). In place of the controller 821, the BB processor 826 may have some or all of the aforementioned logical functions. The BB processor 826 may be a memory that stores communication control programs, or a module including a processor configured to execute programs and associated circuitry. Program updates can modify the functionality of the BB processor 826. This module may be a card or blade inserted into a slot in the base station device 820. Alternatively, the module may be a chip mounted on the card or blade. Meanwhile, the RF circuit 827 may include, for example, a mixer, a filter, and an amplifier, and transmit and receive wireless signals via the antenna 810 .

[0128] As shown in FIG13 , the wireless communication interface 825 may include multiple BB processors 826. For example, multiple BB processors 826 may be compatible with multiple frequency bands used by the eNB 800. As shown in FIG13 , the wireless communication interface 825 may include multiple RF circuits 827. For example, multiple RF circuits 827 may be compatible with multiple antenna elements. Although FIG13 illustrates an example in which the wireless communication interface 825 includes multiple BB processors 826 and multiple RF circuits 827, the wireless communication interface 825 may also include a single BB processor 826 or a single RF circuit 827.

[0129] In the eNB 800 shown in FIG13 , the communication unit 102 and transceiver of the electronic device 100 may be implemented by the wireless communication interface 825. At least a portion of the functionality may also be implemented by the controller 821. For example, the controller 821 may determine appropriate wireless communication terminals participating in FRL by executing the functions of the determination unit 101, the communication unit 102, and the acquisition unit 103, thereby effectively reducing the latency of global model updates and improving the convergence speed of the global model.

[0130] (Second application example)

[0131] FIG14 is a block diagram illustrating a second example of a schematic configuration of an eNB or gNB to which the techniques of this disclosure can be applied. Note that similarly, the following description uses an eNB as an example, but is equally applicable to a gNB. An eNB 830 includes one or more antennas 840, a base station device 850, and an RRH 860. The RRH 860 and each antenna 840 can be connected to each other via an RF cable. The base station device 850 and the RRH 860 can be connected to each other via a high-speed line such as an optical fiber cable.

[0132] Each of the antennas 840 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for RRH 860 to transmit and receive wireless signals. As shown in FIG14 , eNB 830 may include multiple antennas 840. For example, multiple antennas 840 may be compatible with multiple frequency bands used by eNB 830. Although FIG14 shows an example in which eNB 830 includes multiple antennas 840, eNB 830 may also include a single antenna 840.

[0133] Base station device 850 includes a controller 851, a memory 852, a network interface 853, a wireless communication interface 855, and a connection interface 857. Controller 851, memory 852, and network interface 853 are the same as controller 821, memory 822, and network interface 823 described with reference to FIG.

[0134] The wireless communication interface 855 supports any cellular communication scheme (such as LTE and LTE-Advanced) and provides wireless communication to terminals located in the sector corresponding to the RRH 860 via the RRH 860 and the antenna 840. The wireless communication interface 855 may generally include, for example, a BB processor 856. The BB processor 856 is the same as the BB processor 826 described with reference to FIG. 13, except that the BB processor 856 is connected to the RF circuit 864 of the RRH 860 via the connection interface 857. As shown in FIG. 14, the wireless communication interface 855 may include multiple BB processors 856. For example, the multiple BB processors 856 may be compatible with multiple frequency bands used by the eNB 830. Although FIG. 14 shows an example in which the wireless communication interface 855 includes multiple BB processors 856, the wireless communication interface 855 may also include a single BB processor 856.

[0135] The connection interface 857 is an interface for connecting the base station device 850 (wireless communication interface 855) to the RRH 860. The connection interface 857 may also be a communication module for connecting the base station device 850 (wireless communication interface 855) to the RRH 860 for communication in the high-speed line.

[0136] The RRH 860 includes a connection interface 861 and a wireless communication interface 863 .

[0137] The connection interface 861 is an interface for connecting the RRH 860 (wireless communication interface 863) to the base station device 850. The connection interface 861 may also be a communication module for communication in the above-mentioned high-speed line.

[0138] The wireless communication interface 863 transmits and receives wireless signals via the antenna 840. The wireless communication interface 863 may generally include, for example, an RF circuit 864. The RF circuit 864 may include, for example, a mixer, a filter, and an amplifier, and transmits and receives wireless signals via the antenna 840. As shown in FIG14 , the wireless communication interface 863 may include multiple RF circuits 864. For example, the multiple RF circuits 864 may support multiple antenna elements. Although FIG14 shows an example in which the wireless communication interface 863 includes multiple RF circuits 864, the wireless communication interface 863 may also include a single RF circuit 864.

[0139] In the eNB 830 shown in FIG14 , the communication unit 102 and transceiver of the electronic device 100 may be implemented by the wireless communication interface 855 and / or the wireless communication interface 863. At least a portion of the functionality may also be implemented by the controller 851. For example, the controller 851 may determine appropriate wireless communication terminals participating in FRL by executing the functions of the determination unit 101, the communication unit 102, and the acquisition unit 103, thereby effectively reducing the latency of global model updates and improving the convergence speed of the global model.

[0140] [Application examples on user devices]

[0141] (First application example)

[0142] 15 is a block diagram showing an example of a schematic configuration of a smartphone 900 to which the technology of the present disclosure can be applied. The smartphone 900 includes a processor 901, a memory 902, a storage device 903, an external connection interface 904, a camera 906, a sensor 907, a microphone 908, an input device 909, a display device 910, a speaker 911, a wireless communication interface 912, one or more antenna switches 915, one or more antennas 916, a bus 917, a battery 918, and an auxiliary controller 919.

[0143] The processor 901 may be, for example, a CPU or a system on a chip (SoC), and controls the functions of the application layer and other layers of the smartphone 900. The memory 902 includes RAM and ROM, and stores data and programs executed by the processor 901. The storage device 903 may include storage media such as semiconductor memories and hard disks. The external connection interface 904 is an interface for connecting external devices (such as memory cards and universal serial bus (USB) devices) to the smartphone 900.

[0144] The camera 906 includes an image sensor such as a charge coupled device (CCD) and a complementary metal oxide semiconductor (CMOS) and generates a captured image. The sensor 907 may include a group of sensors such as a measurement sensor, a gyroscope sensor, a geomagnetic sensor, and an acceleration sensor. The microphone 908 converts the sound input to the smartphone 900 into an audio signal. The input device 909 includes, for example, a touch sensor, a keypad, a keyboard, a button, or a switch configured to detect a touch on the screen of the display device 910, and receives an operation or information input from the user. The display device 910 includes a screen such as a liquid crystal display (LCD) and an organic light emitting diode (OLED) display and displays an output image of the smartphone 900. The speaker 911 converts the audio signal output from the smartphone 900 into sound.

[0145] The wireless communication interface 912 supports any cellular communication scheme (such as LTE and LTE-Advanced) and performs wireless communications. The wireless communication interface 912 may typically include, for example, a BB processor 913 and an RF circuit 914. The BB processor 913 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and may also perform various types of signal processing for wireless communications. Meanwhile, the RF circuit 914 may include, for example, mixers, filters, and amplifiers, and transmit and receive wireless signals via an antenna 916. Note that while the figure shows a scenario where one RF link is connected to one antenna, this is merely illustrative, and also encompasses scenarios where one RF link is connected to multiple antennas via multiple phase shifters. The wireless communication interface 912 may be a chip module on which the BB processor 913 and RF circuit 914 are integrated. As shown in FIG15 , the wireless communication interface 912 may include multiple BB processors 913 and multiple RF circuits 914. While FIG15 illustrates an example in which the wireless communication interface 912 includes multiple BB processors 913 and multiple RF circuits 914, the wireless communication interface 912 may also include a single BB processor 913 or a single RF circuit 914.

[0146] In addition, in addition to the cellular communication scheme, the wireless communication interface 912 can support other types of wireless communication schemes, such as a short-range wireless communication scheme, a near-field communication scheme, and a wireless local area network (LAN) scheme. In this case, the wireless communication interface 912 may include a BB processor 913 and an RF circuit 914 for each wireless communication scheme.

[0147] Each of the antenna switches 915 switches a connection destination of the antenna 916 between a plurality of circuits (eg, circuits for different wireless communication schemes) included in the wireless communication interface 912 .

[0148] Each of the antennas 916 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals via the wireless communication interface 912. As shown in FIG15 , the smartphone 900 may include multiple antennas 916. Although FIG15 shows an example in which the smartphone 900 includes multiple antennas 916, the smartphone 900 may also include a single antenna 916.

[0149] In addition, the smartphone 900 may include an antenna 916 for each wireless communication scheme. In this case, the antenna switch 915 may be omitted from the configuration of the smartphone 900.

[0150] The bus 917 connects the processor 901, the memory 902, the storage device 903, the external connection interface 904, the camera 906, the sensor 907, the microphone 908, the input device 909, the display device 910, the speaker 911, the wireless communication interface 912, and the auxiliary controller 919. The battery 918 supplies power to the various blocks of the smartphone 900 shown in FIG15 via feeders, which are partially shown as dotted lines in the figure. The auxiliary controller 919 operates the minimum necessary functions of the smartphone 900, for example, in sleep mode.

[0151] In the smartphone 900 shown in FIG15 , the communication unit 202 and the transceiver of the electronic device 200 may be implemented by the wireless communication interface 912. At least a portion of the functionality may also be implemented by the processor 901 or the auxiliary controller 919. For example, the processor 901 or the auxiliary controller 919 may implement the training of the local learning model and the timely updating of the global model by executing the functions of the training unit 201 and the communication unit 202.

[0152] (Second application example)

[0153] 16 is a block diagram showing an example of a schematic configuration of a car navigation device 920 to which the technology of the present disclosure can be applied. The car navigation device 920 includes a processor 921, a memory 922, a global positioning system (GPS) module 924, a sensor 925, a data interface 926, a content player 927, a storage medium interface 928, an input device 929, a display device 930, a speaker 931, a wireless communication interface 933, one or more antenna switches 936, one or more antennas 937, and a battery 938.

[0154] The processor 921 may be, for example, a CPU or an SoC, and controls a navigation function and other functions of the car navigation apparatus 920. The memory 922 includes a RAM and a ROM, and stores data and programs executed by the processor 921.

[0155] The GPS module 924 measures the position (such as latitude, longitude, and altitude) of the car navigation device 920 using GPS signals received from GPS satellites. The sensor 925 may include a group of sensors such as a gyroscope sensor, a geomagnetic sensor, and an air pressure sensor. The data interface 926 is connected to, for example, the in-vehicle network 941 via an unillustrated terminal and acquires data generated by the vehicle (such as vehicle speed data).

[0156] The content player 927 reproduces content stored in a storage medium (such as a CD or DVD) inserted into the storage medium interface 928. The input device 929 includes, for example, a touch sensor, button, or switch configured to detect a touch on the screen of the display device 930, and receives an operation or information input from the user. The display device 930 includes a screen such as an LCD or OLED display and displays an image of a navigation function or reproduced content. The speaker 931 outputs the sound of the navigation function or the reproduced content.

[0157] The wireless communication interface 933 supports any cellular communication scheme (such as LTE and LTE-Advanced) and performs wireless communication. The wireless communication interface 933 may generally include, for example, a BB processor 934 and an RF circuit 935. The BB processor 934 may perform, for example, encoding / decoding, modulation / demodulation, and multiplexing / demultiplexing, and perform various types of signal processing for wireless communication. Meanwhile, the RF circuit 935 may include, for example, a mixer, a filter, and an amplifier, and transmit and receive wireless signals via an antenna 937. The wireless communication interface 933 may also be a chip module on which the BB processor 934 and the RF circuit 935 are integrated. As shown in Figure 16, the wireless communication interface 933 may include multiple BB processors 934 and multiple RF circuits 935. Although Figure 16 shows an example in which the wireless communication interface 933 includes multiple BB processors 934 and multiple RF circuits 935, the wireless communication interface 933 may also include a single BB processor 934 or a single RF circuit 935.

[0158] In addition, in addition to the cellular communication scheme, the wireless communication interface 933 can support other types of wireless communication schemes, such as a short-range wireless communication scheme, a near field communication scheme, and a wireless LAN scheme. In this case, for each wireless communication scheme, the wireless communication interface 933 can include a BB processor 934 and an RF circuit 935.

[0159] Each of the antenna switches 936 switches a connection destination of the antenna 937 between a plurality of circuits included in the wireless communication interface 933 , such as circuits for different wireless communication schemes.

[0160] Each of the antennas 937 includes a single or multiple antenna elements (such as multiple antenna elements included in a MIMO antenna) and is used for transmitting and receiving wireless signals via the wireless communication interface 933. As shown in FIG16, the car navigation device 920 may include multiple antennas 937. Although FIG16 shows an example in which the car navigation device 920 includes multiple antennas 937, the car navigation device 920 may also include a single antenna 937.

[0161] Furthermore, the car navigation device 920 may include an antenna 937 for each wireless communication scheme. In this case, the antenna switch 936 may be omitted from the configuration of the car navigation device 920.

[0162] The battery 938 supplies power to the respective blocks of the car navigation apparatus 920 shown in Fig. 16 via a feeder line, which is partially shown as a dotted line in the figure. The battery 938 accumulates the power supplied from the vehicle.

[0163] In the car navigation device 920 shown in FIG16 , the communication unit 202 and transceiver of the electronic device 200 may be implemented by the wireless communication interface 933. At least a portion of the functionality may also be implemented by the processor 921. For example, the processor 921 may implement the training of the local learning model and the timely updating of the global model by executing the functions of the training unit 201 and the communication unit 202.

[0164] The technology of the present disclosure can also be implemented as an in-vehicle system (or vehicle) 940 including a car navigation device 920, an in-vehicle network 941, and one or more blocks of a vehicle module 942. The vehicle module 942 generates vehicle data (such as vehicle speed, engine speed, and fault information) and outputs the generated data to the in-vehicle network 941.

[0165] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be pointed out that for those skilled in the art, it is understandable that all or any steps or components of the methods and devices of the present disclosure can be implemented in any computing device (including a processor, storage medium, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof. This can be achieved by those skilled in the art using their basic circuit design knowledge or basic programming skills after reading the description of the present disclosure.

[0166] Furthermore, the present disclosure also provides a program product storing machine-readable instruction codes. When the instruction codes are read and executed by a machine, the method according to the embodiment of the present disclosure can be executed.

[0167] Accordingly, the storage medium for carrying the program product storing the machine-readable instruction code is also included in the disclosure of the present invention, including but not limited to a floppy disk, an optical disk, a magneto-optical disk, a memory card, a memory stick, and the like.

[0168] When the present disclosure is implemented through software or firmware, the programs constituting the software are installed from a storage medium or a network to a computer with a dedicated hardware structure (such as the general-purpose computer 1700 shown in Figure 17). When various programs are installed on the computer, it can perform various functions, etc.

[0169] 17 , a central processing unit (CPU) 1701 executes various processes according to a program stored in a read-only memory (ROM) 1702 or a program loaded from a storage section 1708 to a random access memory (RAM) 1703. In the RAM 1703, data required when the CPU 1701 executes various processes, etc., is also stored as needed. The CPU 1701, the ROM 1702, and the RAM 1703 are connected to each other via a bus 1704. An input / output interface 1705 is also connected to the bus 1704.

[0170] The following components are connected to the input / output interface 1705: an input section 1706 (including a keyboard, a mouse, etc.), an output section 1707 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and speakers, etc.), a storage section 1708 (including a hard disk, etc.), and a communication section 1709 (including a network interface card such as a LAN card, a modem, etc.). The communication section 1709 performs communication processing via a network such as the Internet. A drive 1710 may also be connected to the input / output interface 1705 as needed. A removable medium 1711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in the drive 1710 as needed, so that a computer program read therefrom is installed in the storage section 1708 as needed.

[0171] In the case of realizing the above-described series of processing by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 1711 .

[0172] It should be understood by those skilled in the art that such storage media is not limited to the removable medium 1711 shown in FIG. 17 , which stores the program therein and is distributed separately from the device to provide the program to the user. Examples of the removable medium 1711 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memories (CD-ROMs) and digital versatile disks (DVDs)), magneto-optical disks (including minidiscs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be the ROM 1702, a hard disk included in the storage section 1708, or the like, in which the program is stored and distributed to the user together with the device containing them.

[0173] It should also be noted that in the apparatus, method, and system of the present disclosure, each component or step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure. Furthermore, the steps of performing the above series of processes can naturally be performed in chronological order according to the order of description, but do not necessarily need to be performed in chronological order. Certain steps can be performed in parallel or independently of each other.

[0174] Finally, it should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Furthermore, in the absence of further limitations, an element defined by the phrase "comprises a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0175] Although the embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, it should be understood that the embodiments described above are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Those skilled in the art will appreciate that various modifications and variations can be made to the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is solely defined by the appended claims and their equivalents.

Claims

1. An electronic device for wireless communication, comprising: The processing circuit is configured to: Determining wireless communication terminals to participate in federated reinforcement learning based at least on the processing capabilities of wireless communication terminals within coverage of the wireless transceiver node and the wireless communication environment; and Respective local learning models are obtained from wireless communication terminals participating in the federated reinforcement learning, and an updated global model is obtained based on the local learning models.

2. The electronic device according to claim 1, wherein The wireless communication environment of the wireless communication terminal includes one or more of the following: wireless channel quality, data rate, interference intensity, geographical location, information transmission path loss related to the geographical location, and moving speed.

3. The electronic device according to claim 1, wherein The processing circuit is further configured to provide the updated global model to wireless communication terminals within the coverage area.

4. The electronic device according to claim 3, wherein The processing circuit is configured to provide a complete updated global model to wireless communication terminals participating in the federated reinforcement learning, and to provide a lightweight updated global model to wireless communication terminals not participating in the federated reinforcement learning.

5. The electronic device according to claim 3, wherein The processing circuit is configured to provide the updated global model to wireless communication terminals participating in the federated reinforcement learning, wherein wireless communication terminals not participating in the federated reinforcement learning obtain the updated global model from the wireless communication terminals participating in the federated reinforcement learning through P2P communication. The electronic device according to claim 1 , wherein: The processing circuit is configured to obtain information about the wireless communication environment and the processing capability at least in part from the wireless communication terminal.

7. The electronic device according to claim 1, wherein The processing circuit is also configured to schedule a first wireless communication terminal participating in the federated reinforcement learning to assist a second wireless communication terminal participating in the federated reinforcement learning in uploading a local learning model, wherein the second wireless communication terminal sends part or all of the local learning model to the first wireless communication terminal through P2P communication for uploading by the first wireless communication terminal.

8. The electronic device according to claim 7, wherein: The P2P communication includes direct communication link communication.

9. The electronic device according to claim 1, wherein The processing circuit is configured to schedule wireless communication terminals participating in the federated reinforcement learning to upload local learning models, so that the end time of model uploading of each wireless communication terminal remains consistent.

10. The electronic device according to claim 1, wherein The processing circuit is also configured to instruct the wireless communication terminals participating in the federated reinforcement learning to perform model distillation to control the data amount of the local learning model to a predetermined amount.

11. The electronic device according to claim 1, wherein When determining a wireless communication terminal to participate in federated reinforcement learning, the wireless communication environment of the wireless communication terminal has a higher weight than the processing capability of the wireless communication terminal.

12. The electronic device according to claim 1, wherein The wireless communication terminal is a vehicle or a vehicle in a fleet.

13. The electronic device according to claim 12, wherein: The processing capability of the vehicle includes one or more of the following: the vehicle's autonomous driving level, and the vehicle's environmental perception capability.

14. The electronic device according to claim 12, wherein: The vehicle sets a predetermined reward function during training of the local learning model, where the reward function includes one or more of the following: a forward driving reward, a collision avoidance reward, and a speed maintenance reward.

15. An electronic device for wireless communication, comprising: The processing circuit is configured to: In response to confirmation information from the wireless transceiver node, performing training of a local learning model at the wireless communication terminal, wherein the confirmation information indicates that the wireless transceiver node determines that the wireless communication terminal is to participate in federated reinforcement learning based on a processing capability of the wireless communication terminal and a wireless communication environment; and The local learning model is uploaded to the wireless transceiver node and an updated global model is obtained from the wireless transceiver node.

16. The electronic device according to claim 15, wherein The processing circuit is further configured to provide the updated global model to a wireless communication terminal that does not participate in the federated reinforcement learning in response to a request thereof through P2P communication.

17. The electronic device according to claim 15, wherein: The processing circuit is further configured to upload a part or all of the local learning model of the wireless communication terminal with the assistance of other wireless communication terminals participating in the federated reinforcement learning.

18. The electronic device according to claim 17, wherein: The processing circuit is configured to determine the other wireless communication terminal by broadcasting request information, and the wireless communication terminal performs P2P communication with the other wireless communication terminal.

19. The electronic device according to claim 18, wherein The request information includes one or more of the following: an identifier of the wireless communication terminal, a size of a model to be uploaded, and an incentive strategy.

20. The electronic device according to claim 17, wherein The processing circuit is configured to perform model transmission between the wireless communication terminal and the other wireless communication terminals through P2P communication based on the scheduling information from the wireless transceiver node.

21. The electronic device according to claim 15, wherein The wireless communication environment of the wireless communication terminal includes one or more of the following: wireless channel quality, data rate, interference intensity, geographical location, information transmission path loss related to the geographical location, and moving speed.

22. The electronic device according to claim 15, wherein The wireless communication terminal is a vehicle or a vehicle in a fleet.

23. The electronic device according to claim 22, wherein: The processing capability of the vehicle includes one or more of the following: the vehicle's autonomous driving level, and the vehicle's environmental perception capability.

24. The electronic device according to claim 22, wherein The processing circuit is configured to set a predetermined reward function when training the local learning model, where the reward function includes one or more of the following: a forward driving reward, a collision avoidance reward, and a speed maintenance reward.

25. The electronic device according to claim 22, wherein In the case where the wireless communication terminal is a pilot vehicle in a fleet, the processing circuit is further configured to perform a coarse update of the fleet local learning model based on the local learning models of the fleet members, and upload the coarse updated fleet local learning model to the wireless transceiver node.

26. The electronic device according to claim 25, wherein The processing circuit is further configured to send the updated global model obtained from the wireless transceiver node to each team member.

27. A method for wireless communication, comprising: At least based on the processing capability and wireless communication terminal within the coverage of the wireless transceiver node In the online communication environment, determine the wireless communication terminals that will participate in federated reinforcement learning; as well as Respective local learning models are obtained from wireless communication terminals participating in the federated reinforcement learning, and an updated global model is obtained based on the local learning models.

28. A method for wireless communication, comprising: In response to confirmation information from the wireless transceiver node, performing training of a local learning model at the wireless communication terminal, wherein the confirmation information indicates that the wireless transceiver node determines that the wireless communication terminal is to participate in federated reinforcement learning based on a processing capability of the wireless communication terminal and a wireless communication environment; and The local learning model is uploaded to the wireless transceiver node and an updated global model is obtained from the wireless transceiver node.

29. A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to claim 27 or 28.

Citation Information

Cited By

  • Vehicle control method, device, equipment, medium and system based on vehicle key

    CN120544300A