Model training method and communication apparatus

By adjusting the dataset according to the target parameters of the terminal device through network devices, the problem of low model training efficiency of terminal devices is solved, and data transmission is optimized and model convergence is accelerated.

WO2026092307A1PCT designated stage Publication Date: 2026-05-07HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-10-24
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

During the training of artificial intelligence models, the data transmission costs of terminal devices are high, the time consumption is long, and the resources are consumed, resulting in low training efficiency.

Method used

Network devices receive target parameters reported by terminal devices and adjust the amount and difficulty of the dataset to guide model training in real time, including sending datasets and instruction information to optimize the model training process.

Benefits of technology

It improves the efficiency of model training, reduces data transmission overhead, and accelerates the model convergence process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025129716_07052026_PF_FP_ABST
    Figure CN2025129716_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of communications. Provided are a model training method and a communication apparatus, which are conducive to the improvement of model training supporting efficiency. The method may comprise: receiving first indication information, wherein the first indication information is configured to indicate target parameters during the process of training a first model, and the target parameters comprise at least one of the following: one or more results of performance metrics during the process of training the first model, data volumes subsequently required for one or more stages during the process of training the first model, or the training completion progress of the one or more stages during the process of training the first model; and determining the target parameters on the basis of the first indication information. That is, the present application can guide the training of a model on the basis of the target parameters during the model training process.
Need to check novelty before this filing date? Find Prior Art

Description

A model training method and communication device

[0001] This application claims priority to Chinese Patent Application No. 202411540256.1, filed with the State Intellectual Property Office of China on October 30, 2024, entitled “A Model Training Method and Communication Device”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of communication technology, and in particular to a model training method and a communication device. Background Technology

[0003] During the training of artificial intelligence models, network devices send datasets to terminals to support model training on the terminal side. Typically, the amount of data required for model training on the terminal side is large. As a result, the data transmission cost is high, the training time on the terminal side is long, the resources are consumed, and the training efficiency is low. Summary of the Invention

[0004] This application provides a model training method and a communication device, which are beneficial for improving the efficiency of model training.

[0005] Firstly, a model training method is provided. The execution subject of this method can be a network device, a component or device applied to the network device (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the functions of the network device. The method includes: receiving first indication information, the first indication information indicating target parameters in a first model training process, the target parameters including at least one of the following: one or more results of performance metrics in the training of the first model, the amount of data required for subsequent stages of one or more stages in the training of the first model, or the training completion progress of one or more stages in the training of the first model; and determining the target parameters based on the first indication information.

[0006] In this way, based on the target parameters reported by the terminal device during model training to the network device, the network device can understand the terminal device's model training process in a timely manner based on the target parameters. This is beneficial for the network device to guide the terminal device in model training based on the target parameters, thereby improving the efficiency of model training on the terminal device and accelerating model convergence on the terminal device side. In other words, it can avoid the problem of low model training efficiency when the network device does not understand the terminal device's model training process and does not participate in guiding the terminal device's model training.

[0007] In one possible design, the method further includes sending a first dataset, which is determined based on target parameters and is used for training the first model. That is, the network device can determine the first dataset based on first indication information, or in other words, the first dataset is related to the target parameters reported by the terminal device. This is equivalent to the network device determining the next batch of datasets to be sent based on the target parameters during the terminal device's first model training process, enabling real-time adjustment of the dataset based on the training progress of the first model to improve model training efficiency on the terminal device side.

[0008] In one possible design, the first dataset is determined based on target parameters, including the data size and / or data difficulty of the first dataset. In other words, the network device can adjust the data size and / or data difficulty of the first dataset to be sent in a timely manner based on the target parameters reported by the terminal device, to save on data transmission overhead, or to accelerate the convergence of the model trained on the terminal device and improve model training efficiency when the data difficulty is reduced.

[0009] In one possible design, the method further includes sending a second indication message indicating that the first model training is complete or paused. In this case, the network device determines the target performance achieved by the first model based on the target parameters reported by the terminal device, allowing training to be terminated and saving the overhead of transmitting the remaining dataset.

[0010] In one possible design, before receiving the first indication information, the method further includes: sending a second dataset, wherein the first indication information is determined based on the second dataset, which is used for training the first model. That is, when the terminal device trains the first model based on the received second dataset, it simultaneously records the target parameters obtained during training, so that these target parameters can be reported to the network device as part of the first indication information, allowing the network device to adjust the next batch of first datasets to be sent based on the target parameters. Alternatively, the first indication information is related to the second dataset.

[0011] In one possible design, performance metrics include the results of one or more target metrics obtained during the first training phase using a second dataset, where the first training phase comprises one or more training epochs. The target metrics measure the difference between the predicted values ​​and the true values ​​of the first model. The number of training epochs can be predefined or preconfigured. This allows the network device to determine the progress of the currently trained first model based on the target metric results, adjusting the next batch of datasets to be sent and improving model training efficiency on the terminal device side.

[0012] In one possible design, the target metric includes at least one of the following: the value of the loss function of the second dataset, the normalized mean squared error (NMSE), or the cosine similarity.

[0013] In one possible design, the difficulty of the first dataset differs from that of the second dataset. This allows the network device to adjust the difficulty of the dataset to be sent based on the target parameters reported by the terminal device, accelerating model training convergence on the terminal device side and improving training efficiency.

[0014] In one possible design, when the magnitude of change in multiple performance metrics gradually decreases, the data difficulty of the first dataset is greater than that of the second dataset. This gradual decrease in the magnitude of change in multiple performance metrics can be interpreted as a slow decline in the value of the performance metric, such as the loss function. In this case, a more difficult dataset needs to be trained to accelerate model convergence.

[0015] In one possible design, the second dataset serves as the validation set for the first model. That is, this application can also train the model using the validation set to obtain the target parameters.

[0016] In one possible design, when the amount of data required for subsequent stages in the training of the first model decreases, or when the increase in training completion rate for each stage increases, the size of the first dataset is smaller than the size of the second dataset. Conversely, when the amount of data required for subsequent stages in the training of the first model increases, or when the increase in training completion rate for each stage decreases, the size of the first dataset is greater than or equal to the size of the second dataset. This way, when the amount of data required decreases, the overhead of transmitting the dataset can be saved. When the amount of data required increases, the size of the dataset can be increased to accelerate model training convergence.

[0017] In one possible design, before receiving the first indication information, the method further includes sending a third indication information, which is used to indicate the target parameters to be reported during the training of the first model. In this way, when the network device instructs the terminal device to report the target parameters via the third indication information, the target parameters that the terminal device needs to report can be flexibly configured.

[0018] In one possible design, the method further includes: sending an identifier of the training set and / or an identifier of the first dataset; or, sending an identifier of the first model and / or an identifier of the first dataset; or, sending a function identifier corresponding to the first model and / or an identifier of the first dataset. Alternatively, the first dataset can be indicated by at least one of the following: an identifier of the first dataset; an identifier of the first model; or a function identifier corresponding to the first model. This helps the terminal device identify datasets used to train the same model.

[0019] Secondly, a model training method is provided. The execution subject of this method can be a terminal device, a component or device applied to the terminal device (e.g., a processor, chip, or chip system), or a logic module or software capable of implementing all or part of the terminal device's functions. The method includes: determining first indication information, the first indication information being used to indicate target parameters in the first model training process, the target parameters including at least one of the following: one or more results of performance metrics in the first model training process, the amount of data required for subsequent stages of one or more stages in the first model training process, or the training completion progress of one or more stages in the first model training process; and sending the first indication information.

[0020] For the beneficial effects of the second aspect, please refer to the explanation of the first aspect.

[0021] In one possible design, the method further includes: receiving a first dataset, which is determined based on target parameters, and the first dataset is used for training a first model.

[0022] In one possible design, the first dataset is determined based on target parameters, including: the data size and / or data difficulty of the first dataset are determined based on target parameters.

[0023] In one possible design, the method further includes receiving a second indication message indicating that the first model training is complete or training is paused.

[0024] In one possible design, before sending the first instruction information, the method further includes: receiving a second dataset, the second dataset being used for training the first model; determining the first instruction information includes: determining the first instruction information based on the second dataset.

[0025] In one possible design, the performance metrics include the results of one or more target metrics obtained during a first training phase using a second dataset, the first training phase comprising one or more training epochs; wherein the results of the target metrics are used to measure the difference between the predicted values ​​and the true values ​​of the first model.

[0026] In one possible design, the target metric includes at least one of the following: the value of the loss function of the second dataset, the normalized mean square error (NMSE), or the cosine similarity.

[0027] In one possible design, the data difficulty of the first dataset is different from that of the second dataset.

[0028] In one possible design, as the variation in multiple results of the performance metrics gradually decreases, the data difficulty of the first dataset is greater than that of the second dataset.

[0029] In one possible design, the second dataset serves as the validation set for the first model.

[0030] In one possible design, when the amount of data required for subsequent stages in the training of the first model decreases, or the increase in the training completion rate of the multiple stages in the training of the first model increases, the amount of data in the first dataset is less than the amount of data in the second dataset; when the amount of data required for subsequent stages in the training of the first model increases, or the increase in the training completion rate of the multiple stages in the training of the first model decreases, the amount of data in the first dataset is greater than or equal to the amount of data in the second dataset.

[0031] In one possible design, before determining the first indication information, the method further includes: receiving third indication information, which is used to indicate the reporting of target parameters during the training of the first model.

[0032] In one possible design, the method further includes: receiving an identifier of the training set and / or an identifier of the first dataset; or, receiving an identifier of the first model and / or an identifier of the first dataset; or, receiving a function identifier corresponding to the first model and / or an identifier of the first dataset.

[0033] Thirdly, a communication device is provided, which may be a network device. The communication device includes: a transceiver unit for receiving first indication information, the first indication information being used to indicate target parameters in the training process of a first model, the target parameters including at least one of the following: one or more results of performance indicators in the training process of the first model, the amount of data required for subsequent stages of one or more stages in the training process of the first model, or the training completion progress of one or more stages in the training process of the first model; and a processing unit for determining the target parameters based on the first indication information.

[0034] In one possible design, the transceiver unit is also used to send a first dataset, which is determined based on the target parameters, and is used for training the first model.

[0035] In one possible design, the transceiver unit is also used to send a second instruction message, which indicates that the first model training is complete or training is paused.

[0036] In one possible design, the transceiver unit is also used to send a second dataset, the first instruction information being determined based on the second dataset, which is used for training the first model.

[0037] In one possible design, the transceiver unit is also used to send a third instruction message, which is used to instruct the reporting of target parameters during the training of the first model.

[0038] In one possible design, the transceiver unit is also used to send the identifier of the training set and / or the identifier of the first dataset; or, send the identifier of the first model and / or the identifier of the first dataset; or, send the function identifier corresponding to the first model and / or the identifier of the first dataset.

[0039] Fourthly, a communication device is provided, which may be a terminal device. The communication device includes: a processing unit for determining first indication information, the first indication information being used to indicate target parameters in the training process of a first model, the target parameters including at least one of the following: one or more results of performance indicators in the training process of the first model, the amount of data required for subsequent stages of one or more stages in the training process of the first model, or the training completion progress of one or more stages in the training process of the first model; and a transceiver unit for sending the first indication information.

[0040] In one possible design, the transceiver unit is also used to receive a first dataset, which is determined based on the target parameters, and is used for training the first model.

[0041] In one possible design, the transceiver unit is also used to receive a second instruction message, which indicates that the first model training is complete or training is paused.

[0042] In one possible design, the transceiver unit is also used to receive a second dataset, which is used for training the first model; and the processing unit is used to determine the first instruction information based on the second dataset.

[0043] In one possible design, the transceiver unit is also used to receive third instruction information, which is used to instruct the reporting of target parameters during the training of the first model.

[0044] In one possible design, the transceiver unit is also used to receive the identifier of the training set and / or the identifier of the first dataset; or, to receive the identifier of the first model and / or the identifier of the first dataset; or, to receive the function identifier corresponding to the first model and / or the identifier of the first dataset.

[0045] In the third and / or fourth aspects:

[0046] In one possible design, the first dataset is determined based on target parameters, including: the data size and / or data difficulty of the first dataset are determined based on target parameters.

[0047] In one possible design, the performance metrics include the results of one or more target metrics obtained during a first training phase using a second dataset, the first training phase comprising one or more training epochs; wherein the results of the target metrics are used to measure the difference between the predicted values ​​and the true values ​​of the first model.

[0048] In one possible design, the target metric includes at least one of the following: the value of the loss function of the second dataset, the normalized mean squared error (NMSE), or the cosine similarity.

[0049] In one possible design, the data difficulty of the first dataset is different from that of the second dataset.

[0050] In one possible design, as the variation in multiple results of the performance metrics gradually decreases, the data difficulty of the first dataset is greater than that of the second dataset.

[0051] In one possible design, the second dataset serves as the validation set for the first model.

[0052] In one possible design, when the amount of data required for subsequent stages in the training of the first model decreases, or the increase in the training completion rate of the multiple stages in the training of the first model increases, the amount of data in the first dataset is less than the amount of data in the second dataset; when the amount of data required for subsequent stages in the training of the first model increases, or the increase in the training completion rate of the multiple stages in the training of the first model decreases, the amount of data in the first dataset is greater than or equal to the amount of data in the second dataset.

[0053] Fifthly, a method is provided that includes a processor and an interface circuit, the interface circuit being used to receive signals from other communication devices and transmit them to the processor or to send signals from the processor to other communication devices, the processor being used through logic circuits or executing code instructions to implement the method described in the first aspect and any possible design of the first aspect, and / or, the method described in the second aspect and any possible design of the second aspect.

[0054] In a sixth aspect, a computer-readable storage medium is provided, wherein computer instructions are stored therein, which, when executed on a communication device, cause the communication device to perform the method as described in the first aspect and any possible design of the first aspect, and / or, the second aspect and any possible design of the second aspect.

[0055] A seventh aspect provides a computer program product including computer instructions that, when executed on a communication device, cause the communication device to perform the method as described in the first aspect and any possible design of the first aspect, and / or, the second aspect and any possible design of the second aspect.

[0056] Eighthly, a communication system is provided, including a first communication device and a second communication device, wherein the first communication device is configured to perform the method described in the first aspect and any possible design of the first aspect, and the second communication device is configured to perform the method described in the second aspect and any possible design of the second aspect. Attached Figure Description

[0057] Figure 1 is a schematic diagram of two communication systems provided in the embodiments of this application;

[0058] Figure 2 is a schematic diagram of a possible application framework of a communication system provided in an embodiment of this application;

[0059] Figure 3 is a schematic diagram of a possible application framework in a communication system provided in an embodiment of this application;

[0060] Figure 4 is a schematic diagram of batch dataset transmission provided in an embodiment of this application;

[0061] Figure 5 is a flowchart illustrating a model training method provided in an embodiment of this application;

[0062] Figure 6 is a flowchart illustrating a model training method provided in an embodiment of this application;

[0063] Figure 7 is a flowchart illustrating a model training method applied in an O-RAN architecture according to an embodiment of this application;

[0064] Figure 8 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0065] Figure 9 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0066] Figure 10 is a schematic diagram of the structure of a chip provided in an embodiment of this application;

[0067] Figures 11 and 12 are schematic diagrams of neuron structures provided in the embodiments of this application. Detailed Implementation

[0068] The technical solutions provided in this application can be applied to various communication systems, such as 5th generation (5G) or new radio (NR) systems, long term evolution (LTE) systems, LTE frequency division duplex (FDD) systems, LTE time division duplex (TDD) systems, wireless local area network (WLAN) systems, satellite communication systems, and future communication systems, such as integrated systems. The technical solutions provided in this application can also be applied to device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-to-machine (M2M) communication, machine-type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems.

[0069] In a communication system, one network element can send signals to or receive signals from another network element. These signals can include information, signaling, or data. The term "network element" can also be replaced by an entity, network entity, device, communication equipment, communication module, node, communication node, etc. This disclosure uses a network element as an example. For instance, a communication system can include at least one terminal device and at least one network device. The network device can send downlink signals to the terminal device, and / or the terminal device can send uplink signals to the network device. It is understood that the terminal device in this application can be replaced by a first network element, and the network device can be replaced by a second network element, both performing the corresponding communication methods described in this application.

[0070] In wireless communication networks, such as mobile communication networks, the services supported by the networks are becoming increasingly diverse, leading to increasingly diverse requirements. For example, networks need to support ultra-high speeds, ultra-low latency, and / or massive connectivity. This characteristic makes network planning, network configuration, and / or resource scheduling increasingly complex. Furthermore, as network functions become more powerful, such as supporting higher spectrum levels, supporting higher-order multiple-input multiple-output (MIMO) technologies, supporting beamforming, and / or supporting beam management, network energy efficiency has become a hot research topic. These new requirements, new scenarios, and new characteristics bring unprecedented challenges to network planning, operation, and efficient operation. To meet these challenges, AI technology can be introduced into wireless communication networks to achieve network intelligence. To support AI technology in wireless networks, AI nodes may also be introduced.

[0071] Supervised learning, based on collected sample values ​​and labels, uses machine learning algorithms to learn the mapping relationship between sample values ​​and labels, and expresses this learned mapping relationship using a machine learning model. The process of training the machine learning model is the process of learning this mapping relationship. For example, in signal detection, the noisy received signal is the sample, and the corresponding real constellation point is the label. Machine learning aims to learn the mapping relationship between samples and labels through training, that is, to enable the machine learning model to learn a signal detector. During training, the model parameters are optimized by calculating the error between the model's predicted values ​​and the real labels. Once the mapping relationship is learned, it can be used to predict the sample label of each new sample. The mapping relationship learned in supervised learning can include linear mappings and nonlinear mappings. Based on the type of label, the learning task can be divided into classification tasks and regression tasks.

[0072] Unsupervised learning relies solely on collected sample values, using algorithms to discover inherent patterns within the samples. One type of unsupervised learning algorithm uses the samples themselves as supervisory signals; that is, the model learns the mapping relationship from sample to sample, which is called self-supervised learning. During training, model parameters are optimized by calculating the error between the model's predictions and the samples themselves. Self-supervised learning can be used for signal compression and decompression recovery applications; common algorithms include autoencoders and generative adversarial networks.

[0073] Reinforcement learning, unlike supervised learning, is a type of algorithm that learns problem-solving strategies through interaction with the environment. Unlike supervised and unsupervised learning, reinforcement learning problems do not have explicit "correct" action labels. The algorithm needs to interact with the environment to obtain reward signals from the environment, and then adjust its decision actions to obtain a larger reward signal value. For example, in downlink power control, the reinforcement learning model adjusts the downlink transmission power of each user based on the total system throughput feedback from the wireless network, aiming to achieve a higher system throughput. The goal of reinforcement learning is also to learn the mapping relationship between the environment state and the optimal decision action. However, because the label of the "correct action" cannot be obtained in advance, the network cannot be optimized by calculating the error between the action and the "correct action." Reinforcement learning training is achieved through iterative interaction with the environment.

[0074] Deep Neural Networks (DNNs) are a specific implementation of machine learning. According to the general approximation theorem, neural networks can theoretically approximate any continuous function, thus enabling them to learn arbitrary mappings. Traditional communication systems rely on extensive expert knowledge to design communication modules, while DNN-based deep learning communication systems can automatically discover hidden pattern structures from large datasets, establish mapping relationships between data, and achieve performance superior to traditional modeling methods.

[0075] The idea behind DNNs originates from the neuronal structure of the brain. Each neuron performs a weighted summation of its input values, and the result is passed through a non-linear function to generate the output, as shown in Figure 11. Specifically, assume the input to the neuron is x = [x0, ..., x...]. n The weights corresponding to the inputs are w = [w0, ..., w0]. n The bias of the weighted summation is b. The nonlinear function can take many forms; one example is the max{0,x} maximum value function. The effect of a neuron's execution can be... DNNs typically have a multi-layered structure, with each layer containing multiple neurons. The input layer processes the received values ​​through neurons and then passes them to the hidden layers. Similarly, the hidden layers then pass the calculation results to the final output layer, producing the final output of the DNN, as shown in Figure 12.

[0076] DNNs typically have more than one hidden layer, and these hidden layers often directly affect the ability to extract information and fit functions. Increasing the number of hidden layers or widening the width of each layer can improve the function fitting ability of a DNN. The weights in each neuron are the parameters of the DNN network model. The model parameters are optimized through the training process, enabling the DNN network to extract data features and express mapping relationships. DNNs generally use supervised or unsupervised learning strategies to optimize model parameters.

[0077] Based on their construction methods, DNNs can be divided into feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). The image shows an FNN network, characterized by perfectly connected neurons in adjacent layers. This typically requires a large amount of storage space, resulting in high computational complexity.

[0078] CNNs are neural networks specifically designed to process data with a grid-like structure. For example, time-series data (discrete sampling along the time axis) and image data (two-dimensional discrete sampling) can both be considered grid-like data. CNNs do not use all the input information at once for computation; instead, they use a fixed-size window to extract a portion of the information for convolution operations, which significantly reduces the computational cost of model parameters. Furthermore, depending on the type of information extracted by the window (such as people and objects in an image representing different types of information), each window can use different convolution kernels, allowing CNNs to better extract features from the input data.

[0079] Recurrent Neural Networks (RNNs) are a type of distributed neural network (DNN) that utilizes feedback time-series information. Their input includes the current input value and their own output value from the previous time step. RNNs are well-suited for acquiring temporally correlated sequence features, and are particularly applicable to applications such as speech recognition and channel coding / decoding.

[0080] Figure 1(a) is a schematic diagram of a communication system provided in an embodiment of this application. As shown in Figure 1(a), the communication system 100 may include at least one network device, such as network device 110 shown in Figure 1(a); the communication system 100 may also include at least one terminal device, such as terminal device 120 and terminal device 130 shown in Figure 1(a). Network device 110 and terminal devices (such as terminal device 120 and terminal device 130) can communicate via a wireless link. The communication devices in this communication system, for example, network device 110 and terminal device 120, can communicate via multi-antenna technology.

[0081] Figure 1(b) is a schematic diagram of a communication system provided in an embodiment of this application. Compared to the communication system 100 shown in Figure 1(a), the communication system 200 shown in Figure 1(b) further includes an AI network element 140. The AI ​​network element 140 is used to perform AI-related operations, such as building training datasets or training AI models.

[0082] In one possible implementation, network device 110 can send data related to the training of the AI ​​model to AI network element 140, which then constructs a training dataset and trains the AI ​​model. For example, the data related to the training of the AI ​​model may include data reported by the terminal device. AI network element 140 can send the results of operations related to the AI ​​model to network device 110, which then forwards them to the terminal device. For example, the results of operations related to the AI ​​model may include at least one of the following: a trained AI model, model evaluation results, or test results. Exemplarily, a portion of the trained AI model may be deployed on network device 110, and another portion on the terminal device. Alternatively, the trained AI model may be deployed on network device 110. Or, the trained AI model may be deployed on the terminal device.

[0083] It should be understood that Figure 1(b) is only illustrated using the example of AI network element 140 being directly connected to network device 110. In other scenarios, AI network element 140 can also be connected to terminal device. Alternatively, AI network element 140 can be connected to both network device 110 and terminal device simultaneously. Alternatively, AI network element 140 can also be connected to network device 110 through a third-party network element. This application embodiment does not limit the connection relationship between AI network element and other network elements.

[0084] AI element 140 can also be set as a module in network devices and / or terminal devices, for example, in network device 110 or terminal device shown in Figure 1(a).

[0085] Figures 1(a) and 1(b) are simplified schematic diagrams for ease of understanding. For example, the communication system may also include other devices, such as wireless relay devices and / or wireless backhaul devices, which are not shown in Figures 1(a) and 1(b). In practical applications, the communication system may include multiple network devices or multiple terminal devices. This application does not limit the number of network devices and terminal devices included in the communication system.

[0086] In the embodiments of this application, the terminal device may also be referred to as user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user apparatus.

[0087] Terminal devices can be devices that provide voice / data, such as handheld devices with wireless connectivity, in-vehicle devices, etc. Currently, examples of terminals include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, wearable devices, terminal devices in 5G networks, or future public land mobile communication networks. Terminal devices in a network (PLMN), etc., are not limited to this in the embodiments of this application.

[0088] By way of example and not limitation, in this embodiment, the terminal device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0089] In this embodiment, the device for implementing the functions of the terminal device can be the terminal device itself, or it can be any device capable of supporting the terminal device in implementing those functions, such as a chip system. This device can be installed in or used in conjunction with the terminal device. In this embodiment, the chip system can be composed of chips or may include chips and other discrete components. This embodiment only uses the terminal device as an example to illustrate the device for implementing the functions of the terminal device, and does not constitute a limitation on the solution of this embodiment.

[0090] The network device in this application embodiment can be a device for communicating with a terminal device. This network device can also be called an access network device or a wireless access network device, such as a base station. In this application embodiment, the network device can refer to a radio access network (RAN) node (or device) that connects the terminal device to the wireless network. A base station can broadly encompass, or be replaced by, various names including: NodeB, evolved NodeB (eNB), next-generation NodeB (gNB), relay station, access point, transmitting and receiving point (TRP), transmitting point (TP), master station, auxiliary station, motor slide retainer (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), positioning node, etc. A base station can be a macro base station, micro base station, relay node, donor node, or similar entities, or combinations thereof. A base station can also refer to a communication module, modem, or chip installed within the aforementioned equipment or apparatus. A base station can also be a mobile switching center, a device that performs base station functions in D2D, V2X, and M2M communications, or a device that performs base station functions in future communication systems. A base station can support networks using the same or different access technologies. Optionally, a RAN node can also be a server, wearable device, vehicle, or in-vehicle equipment. For example, the access network equipment in vehicle-to-everything (V2X) technology can be a roadside unit (RSU). The embodiments of this application do not limit the specific technologies or equipment forms used in the network equipment.

[0091] Base stations can be fixed or mobile. For example, a helicopter or drone can be configured to act as a mobile base station, and one or more cells can move depending on the location of the mobile base station. In other examples, a helicopter or drone can be configured as a device to communicate with another base station.

[0092] In some deployments, the network devices mentioned in the embodiments of this application may be devices including CU, DU, or CU and DU, or devices with control plane CU nodes (central unit-control plane (CU-CP)) and user plane CU nodes (central unit-user plane (CU-UP)) and DU nodes. For example, the network devices may include gNB-CU-CP, gNB-CU-UP, and gNB-DU.

[0093] In some deployments, multiple RAN nodes collaborate to assist terminals in achieving wireless access, with different RAN nodes each implementing some of the base station's functions. For example, RAN nodes can be CUs, DUs, CU-CPs, CU-UPs, or RUs. CUs and DUs can be configured separately or included in the same network element, such as a BBU. RUs can be included in radio frequency equipment or radio frequency units, such as RRUs, AAUs, or RRHs.

[0094] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an ORAN system, CU can also be called O-CU (open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.

[0095] In this embodiment, the apparatus for implementing the functions of a network device can be a network device itself; it can also be an apparatus capable of supporting the network device in implementing those functions, such as a chip system, hardware circuit, software module, or a hardware circuit plus a software module. This apparatus can be installed in the network device or used in conjunction with the network device. In this embodiment, the example of a network device being used to implement the functions of a network device is provided only and does not constitute a limitation on the solutions described in this embodiment.

[0096] Network devices and / or terminal devices can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they can also be deployed on water; and they can also be deployed in the air on airplanes, balloons, and satellites. This application does not limit the scenario in which the network devices and terminal devices are located. Furthermore, terminal devices and network devices can be hardware devices, or software functions running on dedicated hardware or general-purpose hardware, such as virtualization functions instantiated on a platform (e.g., a cloud platform), or entities that include dedicated or general-purpose hardware devices and software functions. This application does not limit the specific form of the terminal devices and network devices.

[0097] In this embodiment, the AI ​​node can be deployed in one or more of the following locations within the communication system: access network devices, terminal devices, or core network devices, etc. Alternatively, the AI ​​node can be deployed independently, for example, in a location other than any of the aforementioned devices, such as in the host or cloud server of an over-the-top (OTT) system. The AI ​​node can communicate with other devices in the communication system, which can be one or more of the following: network devices, terminal devices, or network elements of the core network, etc.

[0098] It is understood that this application does not limit the number of AI nodes. For example, when there are multiple AI nodes, these nodes can be divided based on function, such as different AI nodes being responsible for different functions.

[0099] It can also be understood that AI nodes can be independent devices, or they can be integrated into the same device to achieve different functions. Alternatively, they can be network elements in hardware devices, software functions running on dedicated hardware, or virtualization functions instantiated on a platform (e.g., a cloud platform). This application does not limit the specific form of the aforementioned AI nodes.

[0100] In this embodiment, the AI ​​node can be an AI network element or an AI module.

[0101] Figure 2 illustrates a possible application framework in a communication system. As shown in Figure 2, network elements in the communication system are connected via interfaces (e.g., NG, Xn) or air interfaces. These network element nodes, such as core network equipment, access network nodes (RAN nodes), terminals, or one or more devices in the OAM, are equipped with one or more AI modules (only one is shown in Figure 2 for clarity). An access network node can be a single RAN node or can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be equipped with one or more AI modules. Optionally, the CU can be further divided into CU-CP and CU-UP. One or more AI models are configured in the CU-CP and / or CU-UP.

[0102] The AI ​​module is used to implement corresponding AI functions. AI modules deployed in different network elements can be the same or different. Depending on the parameter configuration, the AI ​​module can achieve different functions. The AI ​​module model can be configured based on one or more of the following parameters: structural parameters (e.g., at least one of the following: number of neural network layers, neural network width, inter-layer connections, neuron weights, neuron activation function, or bias in the activation function), input parameters (e.g., type and / or dimension of input parameters), or output parameters (e.g., type and / or dimension of output parameters). The bias in the activation function can also be referred to as the neural network bias.

[0103] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or they can be deployed on the same node or device.

[0104] Figure 3 illustrates a possible application framework in a communication system. As shown in Figure 3, the communication system includes a RAN intelligent controller (RIC). For example, the RIC can be an AI module in the core network device or access network node shown in Figure 2, used to implement AI-related functions. The RIC includes near-real-time RICs (near-RT RICs) and non-real-time RICs (non-RT RICs). Non-real-time RICs primarily process non-real-time information, such as data that is not sensitive to latency, with latency in the order of seconds. Real-time RICs primarily process near-real-time information, such as data that is relatively sensitive to latency, with latency in the order of tens of milliseconds.

[0105] In this embodiment, the network device can be a network device equipped with one or more AI modules. The network device can be one or more devices in the core network device, access network node (RAN node), or OAM as shown in Figure 2. For example, the AI ​​module can be the RIC shown in Figure 3, such as a near real-time RIC or a non-real-time RIC. For example, the near real-time RIC is set in the RAN node (e.g., in CU, DU), while the non-real-time RIC is set in the OAM, cloud server, core network device, or other network device. The RIC can obtain a subset from multiple terminal devices from the RAN node (e.g., CU, CU-CP, CU-UP, DU, and / or RU), reassemble it into a training dataset #2, and train based on the training dataset #2. Exemplarily, the near real-time RIC and the non-real-time RIC can also be set as separate network elements, and the network device can be a near real-time RIC or a non-real-time RIC.

[0106] Near real-time RICs are used for model training and inference. For example, they are used to train AI models and then use those models for inference. Non-real-time RICs are also used for model training and inference. For example, they are used to train AI models and then use those models for inference. Near real-time and non-real-time RICs can also be configured as separate network elements. Optionally, near real-time and non-real-time RICs can also be part of other devices. For example, near real-time RICs can be set up in RAN nodes (e.g., CU, DU), while non-real-time RICs can be set up in OAM, cloud servers, core network devices, or other network devices.

[0107] For ease of understanding, the following exemplary descriptions of some concepts related to the embodiments of this application are provided for reference, as shown below.

[0108] AI: To give machines human-like intelligence, using computer hardware and software to simulate certain intelligent human behaviors, including machine learning and many other methods.

[0109] Machine learning (ML) involves learning models or rules from raw data. There are many different machine learning methods, such as neural networks, decision trees, and support vector machines. ML is an important technical approach to AI. Machine learning includes supervised learning, unsupervised learning, and reinforcement learning.

[0110] AI model: This refers to a function model that maps an input of a certain dimension to an output of a certain dimension, and its parameters are obtained through machine learning training. For example, f(x) = ax 2 +b is a quadratic function model, which can be viewed as an AI model. a and b correspond to the parameters of the model, which can be obtained through machine learning training.

[0111] Neural network: Here it refers to artificial neural network, which is a mathematical model that imitates the behavioral characteristics of animal neural networks to perform distributed parallel information processing. It is a special form of AI model.

[0112] Deep neural networks (DNNs) are neural networks with multiple hidden layers. Depending on their construction, DNNs include feedforward neural networks (FNNs), convolutional neural networks (CNNs), or recurrent neural networks (RNNs). These network structures are all built around neurons. Each neuron performs a weighted summation of its input values, and the result is passed through a non-linear function to produce the output. The weights of the weighted summation operation and the non-linear function are called the parameters of the neural network. Taking a neuron with a non-linear function max{0,x} as an example... The parameters of the operated neuron are weights w = [w0, ..., w n The weighted summation bias is b, and the nonlinear function is max{0,x}. The parameters of all neurons in a neural network constitute the parameters of that neural network.

[0113] Deep learning: Machine learning that utilizes deep neural networks.

[0114] Dataset: The data used for model training, validation, and testing in machine learning. The quantity and quality of the data will affect the effectiveness of machine learning.

[0115] Model training: By selecting an appropriate loss function, the model parameters are trained using optimization algorithms to minimize the loss function value.

[0116] Training epoch: The process by which the model performs a complete forward and backward propagation on the training dataset.

[0117] Training batch: The process of taking a small portion of data from the training dataset and propagating this portion of data forward and backward during each iteration of model updates.

[0118] Loss function: Used to measure the difference between the model's predictions and the actual values.

[0119] Model testing: Evaluate model performance using test data after training.

[0120] Model application: Using the trained model to solve practical problems.

[0121] Course-based learning is a training strategy that mimics the learning sequence in human education. Its core idea is to train machine learning models starting with easier data and gradually adding more complex data until the model can be trained on a complete dataset. This method improves the model's generalization ability and convergence speed. Course-based learning mainly consists of two modules: a difficulty measurer and a training scheduler. The difficulty measurer is responsible for measuring the difficulty of each data sample, while the training scheduler is responsible for controlling the difficulty of the data used during training.

[0122] During dual-end or single-end model training between network devices and terminal devices, the network device needs to send a dataset to the terminal device to support model training. Model training can be performed on the terminal device itself or on a non-3GPP entity. If model training is performed on a terminal device, the network device needs to send the complete dataset to support model training based on the complete dataset. If model training is performed on a non-3GPP entity, the network device does not need to send the complete dataset to a single terminal device; instead, it sends different subsets of the dataset to different terminal devices. The different terminal devices then send the received subsets to the non-3GPP entity, which reassembles the subsets to reconstruct the original dataset and performs model training based on it. Each subset is associated with a dataset identifier (ID) to support dataset reconstruction.

[0123] Figure 4 illustrates a batch-based dataset transmission method. The communication system includes K gNBs (gNB#1 to gNB#K) and N UEs and non-3GPP entities. Each of the K gNBs (gNB#1 to gNB#K) transmits a subset of dataset #1 to one of the N terminals. The subset received by each terminal is denoted as subset #N (any one of subset #1, subset #2, ..., subset #N). The subset transmitted by the Kth gNB (gNB#K) to the Nth terminal is denoted as subset #K*N. Thus, the subsets of dataset #1 received by the non-3GPP entities from the N terminals are {subset #1, subset #2, ..., subset #N, ..., subset #K*N}. The non-3GPP entities can reconstruct dataset #1 based on the identifiers of dataset #1 carried in these subsets, and then use dataset #1 for model training.

[0124] Therefore, if model training is performed on a terminal device, the transmission cost of the dataset is high and the resources consumed are significant when a single terminal device receives the complete dataset. The terminal device needs to train on the entire dataset, resulting in a long training time and a high failure rate. Similarly, if model training is performed on a non-3GPP entity, the non-3GPP entity still needs to receive a subset of the dataset from N terminal devices and train the model based on the reconstructed dataset. This also presents the problems of high dataset transmission cost, long training time, high resource consumption, and a high failure rate.

[0125] Therefore, embodiments of this application provide a model training method and a communication device. When the communication device executes the method, the terminal device can report target parameters during the model training process to the network device. For example, the target parameters may be at least one of the following: one or more results of performance metrics during the training of the first model, the amount of data required for subsequent stages of one or more stages in the model training process, or the training completion progress of one or more stages in the model training process. These target parameters are equivalent to auxiliary information reported by the terminal device to the network device during model training. For the network device, timely understanding of the terminal device's model training process based on the target parameters is beneficial for the network device to send the training model dataset to the terminal device based on the target parameters, or to guide the terminal device in model training based on the target parameters, thereby improving the model training efficiency on the terminal device side. This avoids the problem of the network device sending a default full dataset to the terminal device without understanding the terminal device's model training process, leading to low efficiency in model training.

[0126] Figure 5 shows a flowchart of a model training method provided in an embodiment of this application. In this method, the terminal device sends the target parameters during the model training process to the network device through first instruction information, enabling the network device to understand the model training process of the terminal device in a timely manner and guide the terminal device in model training. The method includes the following steps.

[0127] 501. The terminal device sends a first instruction message, which is used to indicate the target parameters in the training process of the first model. The target parameters include at least one of the following: one or more results of performance indicators in the training process of the first model, the amount of data required for the subsequent stages of one or more stages in the training process of the first model, or the training completion progress of one or more stages in the training process of the first model.

[0128] Accordingly, the network device receives the first instruction information.

[0129] In some embodiments, the first indication information used to indicate the target parameters during the training process of the first model includes: the first indication information includes the target parameters during the training process of the first model.

[0130] For example, the target parameters including one or more performance metrics during the training of the first model can be understood as the terminal device obtaining one or more results corresponding to at least one performance metric during the training of the first model. For example, when there is only one performance metric, the one or more results corresponding to that performance metric can be understood as the results of one or more target metrics obtained by the terminal device during the training of the first model.

[0131] For example, the target parameters, including the amount of data required for each of the one or more stages in the training of the first model, can be understood as follows: during the training of the first model, the terminal device can obtain the amount of data required for each training epoch interval in one or more training epochs. For example, it can record the amount of data required for each of the two training epochs at the end of each epoch. After obtaining the amount of data recorded for multiple training epochs, the terminal device can report the recorded amount of data to the network device. Each of the multiple stages here can be understood as each training epoch interval.

[0132] For example, the target parameters, including the training completion progress of one or more stages in the training of the first model, can be understood as follows: during the training of the first model, the terminal device can obtain the training completion progress at intervals between one or more training epochs. For example, it can record the training completion rate of these two training epochs at the end of each epoch. After obtaining the training completion progress recorded for multiple training epochs, the terminal device can report the recorded training completion progress to the network device.

[0133] In some embodiments, the target parameter can also be understood as a process indicator. Process indicators can be used to evaluate the internal operational efficiency of the system's output. Process indicators can also be understood as evaluation indicators.

[0134] 502. The network device determines the target parameters based on the first indication information.

[0135] The network device determines the target parameters based on the first indication information, which can be understood as the network device determining one or more results of the above-mentioned performance indicators, the amount of data required for one or more subsequent stages, or the training completion progress of one or more stages based on the first indication information.

[0136] In this way, for network devices, after determining the target parameters based on the first indication information reported by the terminal device, the network device can promptly understand the progress or status of the first model training on the terminal device side based on the target parameters. This is beneficial for the network device to guide the terminal device side to train the model based on the progress or status, thereby improving the model training efficiency on the terminal device side.

[0137] For example, when determining the target parameters of the terminal device during the model training process, the network device can avoid directly sending the default full dataset to the terminal device or directly sending the full dataset to a non-3GPP entity through the terminal device. Instead, it can send the dataset to the terminal device when it understands the progress of the first model training on the terminal device side, thereby improving the model training efficiency on the terminal device side.

[0138] Therefore, Figure 6 shows a flowchart of a model training method provided in this application embodiment. In this method, the terminal device can determine the target parameters for training the first model based on the dataset received from the network device, and report the target parameters to the network device. The network device adjusts the dataset to be sent next based on the target parameters and sends the dataset adapted to the target parameters to the terminal device.

[0139] For example, taking the network device as the base station and the terminal device as the UE, this method can be applied to the dual-end model interfacing stage between the base station and the UE, the fine-tuning stage of the dual-end model between the base station and the UE, or the single-end model training stage on the UE side. That is, when the base station sends the dataset to the UE, the dataset can be sent in batches. During the model training process, the UE reports relevant target parameters, and the base station can send the next batch of datasets based on the model training status on the terminal side. The following process illustrates the method using the example of the first model training occurring on the UE side, and includes the following steps.

[0140] 601. The base station and UE complete the pairing of the dual-end model.

[0141] The pairing of the base station and UE into a dual-end model can be understood as follows: before the base station and UE use the first model for signaling or data transmission, they interact to inform each other that they need to perform an iterative training process to obtain a dual-end model, thus completing the pairing before the dual-end model training begins. For example, this pairing process includes the pairing of encryption and decryption (enc-dec) during data transmission between the base station and UE.

[0142] For example, this dual-end model is used for the transmission of channel state information (CSI). Before starting dual-end model training, the base station informs the UE that it will begin training the model for transmitting CSI. When the UE decides to train the model for transmitting CSI, it sends a confirmation message to the base station. The model for transmitting CSI can be understood as follows: the UE compresses the CSI matrix based on the model trained on the UE side to obtain a compressed CSI vector, and sends this CSI vector to the base station. The base station decompresses the CSI vector based on the received CSI vector and the model trained on the base station side to obtain the CSI matrix. For example, the dual-end model here is the first model in this application.

[0143] Alternatively, if UE single-end model training is being performed, step 601 can be replaced by the UE reporting the functional ID of the first model to the base station to indicate the type of the first model.

[0144] For example, this function identifier can indicate that the capability of the first model is positioning, meaning the UE can perform positioning based on the first model. The input of the first model is, for example, CSI, and the output is the UE's positioning information, such as latitude and longitude information. Alternatively, this function identifier can indicate that the capability of the first model is beam management, meaning the UE can perform beam management based on the first model. The input of the first model is, for example, CSI, and the output is the UE's beam direction or beam identifier, etc.

[0145] 602. The base station sends a third instruction message, which is used to instruct the reporting of target parameters during the training of the first model.

[0146] Accordingly, the UE receives the third indication information.

[0147] For example, the base station sends third indication information to the UE via the downlink shared channel. The third indication information includes the target parameters that the UE needs to report during the training of the first model.

[0148] In some embodiments, the target parameter includes at least one of the following: one or more results of performance metrics during the training of the first model, the amount of data required for subsequent stages of one or more stages during the training of the first model, or the training completion progress of one or more stages during the training of the first model.

[0149] In some embodiments, the performance metrics include the results of one or more target metrics obtained during a first training phase using a second dataset, wherein the first training phase includes one or more training epochs. The results of the target metrics are used to measure the difference between the predicted values ​​and the true values ​​of the first model.

[0150] The number of training rounds can be predefined or preconfigured. Essentially, the base station instructs the UE to send the results of one or more target metrics obtained by the UE to the base station after performing the first training phase using the received dataset. The second dataset can be understood as any dataset sent by the base station to the UE for training the first model. The second dataset can also be understood as any subset of the training set sent by the UE to the base station for training the first model.

[0151] In some embodiments, the target metric includes at least one of the following: the value of the loss function of the second dataset, the normalized mean square error (NMSE), or the cosine similarity.

[0152] For example, when the performance metric includes the values ​​of loss functions from multiple second datasets obtained during the first training phase using the second dataset, it can be understood that the base station instructs the UE to report curve information showing the changes in the values ​​of the loss functions during the first training phase based on the second datasets. This curve information includes the values ​​of multiple loss functions, allowing the base station to understand the progress of the UE's training of the first model in a timely manner based on the curve information showing the changes in the values ​​of the loss functions. Similarly, when the performance metric includes the NMSE obtained during the first training phase using the second dataset, it can be understood that the base station instructs the UE to report curve information showing the changes in the NMSE during the first training phase using the second dataset. This curve information includes the values ​​of multiple NMSEs. Likewise, when the performance metric includes the cosine similarity obtained during the first training phase using the second dataset, it can be understood that the base station instructs the UE to report curve information showing the changes in the cosine similarity during the first training phase using the second dataset. This curve information includes the values ​​of multiple cosine similarities.

[0153] Among them, the value of the loss function, NMSE, or cosine similarity are all used to measure the difference between the predicted value and the true value of the first model.

[0154] 603. The base station sends a second dataset, which is used to train the first model.

[0155] Accordingly, the UE receives the second dataset.

[0156] In some embodiments, if the first model is a dual-end model, and before the base station and UE have used the first model for data transmission, when the base station sends the second dataset to the UE, the base station may also send the identifier of the training set and / or the identifier of the second dataset to the UE. The training set refers to the full dataset stored by the base station for training the first model on the UE side, including the second dataset and the first dataset in this application.

[0157] For example, the identifier for the training set is denoted as dataset ID, and the identifier for the second dataset is denoted as subset ID. The identifier for the training set and / or the identifier for the second dataset can be sent to the UE in the same signaling message as the identifier for the second dataset, or they can be sent to the UE in different signaling messages. The identifier for the training set can be sent by the base station when sending the identifier for each dataset, or it can be sent to the UE once before sending the first dataset. In this way, for the UE, the dataset used to train the first model can be determined based on the identifiers for the training set and the second dataset, so that the first model can be trained based on the dataset used to train the first model.

[0158] In some embodiments, if the first model is a dual-end model, and the base station and UE have already used the first model for data transmission, and the base station needs to fine-tune the first model, the base station continues to send a dataset to the UE to support the UE in continuing to train the first model based on the received dataset, thereby obtaining a fine-tuned first model. In this case, when the base station sends a second dataset to the UE, considering that the UE may be using multiple models, the base station may send the identifier of the first model and / or the identifier of the first dataset to the UE.

[0159] For example, the identifier for the first model is denoted as model ID, and the identifier for the second dataset is denoted as subset ID. Similarly, the identifier for the first model and / or the identifier for the second dataset can be sent to the UE in the same signaling message as the identifier for the second dataset, or they can be sent to the UE in different signaling messages. The identifier for the first model can be sent by the base station when sending the identifier for each dataset, or it can be sent to the UE once before sending the first dataset. In this way, the UE can determine the dataset used to train the first model based on the identifier for the first model and the identifier for the second dataset, and then perform fine-tuning training of the first model based on the dataset used to train the first model.

[0160] Alternatively, when different model identifiers correspond to different training set identifiers, when the base station instructs the UE to perform fine-tuning training of the first model, the identifier of the first model can be replaced with the identifier of the training set.

[0161] In some embodiments, if the first model is a single-ended model on the UE side, when the base station sends the second dataset to the UE, the base station may also send the function identifier corresponding to the first model and / or the identifier of the second dataset to the UE. That is, the function identifiers corresponding to different models can be used to indicate models with different functions. For example, the function of the first model may be positioning or beam management, etc.

[0162] For example, the functional identifier corresponding to the first model is denoted as the functionality ID, and the identifier of the second dataset is denoted as the subset ID. Similarly, the functional identifier corresponding to the first model and / or the identifier of the second dataset can be sent to the UE in the same signaling as the second dataset, or they can be sent to the UE in different signaling. The functional identifier corresponding to the first model can be sent by the base station when sending the identifier of each dataset, or it can be sent to the UE once before sending the first dataset. In this way, for the UE, the dataset used to train the first model can be determined based on the functional identifier corresponding to the first model and the identifier of the second dataset, so that a single-end model on the UE side, i.e., the first model, can be built based on the dataset used to train the first model. Of course, when fine-tuning the first model on the UE side, the base station can also send the functional identifier corresponding to the first model and / or the identifier of the second dataset simultaneously when sending the dataset to the UE.

[0163] Alternatively, it can be said that the second dataset is indicated by at least one of the following: the identifier of the second dataset; the identifier of the first model; and the functional identifier corresponding to the first model.

[0164] 604. The UE determines the first indication information based on the second dataset. The first indication information is used to indicate the target parameters during the training process of the first model.

[0165] When the UE receives the second dataset, it can train the first model based on the second dataset. The UE can obtain the target parameters during the training process through the following methods: Method 1, Method 2, or Method 3.

[0166] Method 1: If the aforementioned third indication information indicates that the target parameters to be reported by the UE during the training of the first model include one or more results of performance metrics during the training of the first model, and the performance metrics include the results of one or more target metrics obtained during a preset training epoch using the second dataset, the UE can perform multiple training epochs based on the second dataset and record the results of one or more target metrics at the training epoch intervals to obtain the first indication information. The first indication information is used to indicate the results of one or more target metrics at the training epoch intervals. Alternatively, the first indication information is determined based on the second dataset.

[0167] For example, with 10 training epochs and a epoch interval of 2, the UE records the result of the target metric after the second training epoch based on the second dataset, such as the value of the loss function. Similarly, the UE records the result of the target metric after the 4th, 6th, 8th, or 10th training epochs. Thus, after the UE completes 10 training epochs, it obtains the results of 5 target metrics, such as the values ​​of 5 loss functions. That is, the results of the multiple target metrics are the values ​​of 5 loss functions. These 5 loss function values ​​can be used to indicate the change curve of the training set loss function over the 10 training epoch intervals.

[0168] In some embodiments, the second dataset serves as a validation set for the first model. That is, the UE can also train the first model based on the validation set.

[0169] For example, when the base station sends third indication information to the UE, the third indication information also includes indication information for the verification set. Alternatively, the verification set can be carried and sent to the UE in different signaling messages than the third indication information. If the UE trains the first model based on the verification set, the UE can perform one or more training rounds based on the received second dataset and obtain the results of multiple target metrics at training round intervals or training batch intervals. If the UE obtains the results of multiple target metrics at training round intervals, the implementation is similar to the example above. If the results of multiple target metrics at training batch intervals are obtained, it can be understood that one training round includes multiple training batches, and multiple training batches can be executed in each training round. Each training batch can be trained using a portion of the data in the second dataset, and the results of the target metrics obtained at each training batch interval are recorded. In this way, when all the data in the second dataset is trained in multiple batches, the results of the target metrics at multiple training batch intervals in one training round can be obtained. Thus, after performing multiple training rounds based on the second dataset, the results of the target metrics at multiple training batch intervals corresponding to multiple training rounds can be obtained.

[0170] The value of the loss function mentioned above can also be the value of NMSE or the value of cosine similarity. That is, NMSE or cosine similarity can also be understood as a loss function. When the loss function is NMSE, the curve of the change of the value of the loss function can also be recorded as [NMSE1, NMSE2, NMSE3, ...].

[0171] Method 2: If the third indication information indicates that the target parameters to be reported by the UE during the training of the first model include the amount of data required for one or more subsequent stages of the training of the first model, the UE can perform training for one or more stages based on the second dataset and record the amount of data required for one or more subsequent stages to obtain the first indication information. The first indication information is used to indicate the amount of data required for one or more subsequent stages.

[0172] For example, the UE performs 10 training epochs based on the second dataset. These 10 training epochs include stages corresponding to 2, 4, 8, and 10 training epochs respectively. That is, after completing 2 training epochs based on the second dataset, the UE can obtain the amount of data required for subsequent training. Similarly, after completing 4, 8, and 10 training epochs respectively, the UE can obtain the amount of data required for subsequent training. Thus, the target parameters to be reported include the amount of data required for each subsequent stage in the training of the first model, which is equivalent to the target parameters to be reported including the amount of data required for each of the 5 subsequent stages in the training of the first model.

[0173] Method 3: If the third indication information indicates that the target parameters to be reported by the UE during the training of the first model include the training completion progress of one or more stages in the training of the first model, the UE can perform training of one or more stages based on the second dataset and record the training completion degree of one or more stages to obtain the first indication information. The first indication information is used to indicate the training completion degree of one or more stages.

[0174] For example, similar to the example above, the target parameter could include the amount of data required for each of the five stages in training the first model. Alternatively, the target parameter could include the training completion rate of each of the five stages in training the first model. For example, the training completion rate could be indicated as a percentage of training completion.

[0175] 605. The UE sends the first instruction information to the base station.

[0176] Accordingly, the base station receives the first indication information sent by the UE.

[0177] 606. The base station determines the target parameters based on the first indication information, and determines the first dataset based on the target parameters.

[0178] Alternatively, the first dataset is determined based on the target parameters indicated by the first instruction information, and it is used to train the first model. The first dataset can be understood as the next batch of datasets distributed by the base station after the second dataset, and it is related to the target parameters.

[0179] If the UE has multiple models to train, including the first model, and all models are trained on the same dataset, the first instruction information can also indicate the target parameters that the UE needs to report during the training of the second model.

[0180] In some embodiments, the first dataset is determined based on target parameters, including: the data size and / or data difficulty of the first dataset are determined based on target parameters. In other words, the data size and / or data difficulty of the first dataset are related to the target parameters.

[0181] In other words, when the base station receives the first indication information reported by the UE, it can adjust the amount of data and / or difficulty information of the first dataset to be sent in the next batch based on the target parameters.

[0182] For example, the base station determines the difficulty information or data volume of the first dataset based on a pre-configured course learning algorithm and the target parameters indicated by the first indication information, and obtains the first dataset to be transmitted. The course learning algorithm can be understood as a model configured on the base station side, whose input is the target parameters indicated by the first indication information, and whose output is the first dataset.

[0183] In this scenario, the data difficulty of the first dataset differs from that of the second dataset. Alternatively, the data difficulty of the first dataset and the second dataset are the same. Alternatively, the data size of the first dataset is greater than or equal to that of the second dataset. Alternatively, the data size of the first dataset is less than that of the second dataset.

[0184] In some embodiments, based on Method 1, if the target parameters indicated by the first indication information include one or more results of performance metrics during the training of the first model, and the change in one or more results of the performance metrics gradually decreases, the data difficulty of the first dataset is greater than that of the second dataset. Here, one or more results can also be understood as the values ​​or levels of the performance metrics.

[0185] For example, when one or more results of the performance metric are the values ​​of the loss function for the second dataset, it is equivalent to the base station determining the difficulty of the first dataset to be transmitted in the next batch based on the curve of the loss function, and thus obtaining the first dataset. Table 1 shows an example of the relationship between the loss function and the adjusted first dataset.

[0186] Table 1

[0187] For example, when the curve of the loss function mentioned above is shown as [0.83, 0.67, 0.56, 0.54, 0.53], the value of this curve can be understood as the value of the loss function obtained at the interval of the 5 training rounds shown in the example above. Referring to Table 1, if the base station determines the change characteristics of the loss function based on the curve of the loss function, for example, if the change characteristics of the loss function are slow decreases, considering that the loss function is used to measure the difference between the model's predicted value and the true value, if the difference is small, it can be understood that the base station determines that the UE needs to learn a more difficult first dataset to continue training the first model, so that the UE can process the more difficult data sent by the base station based on the first model. If the base station determines that the trend of the change in the value of the loss function is unstable, that is, there is a trend of increasing and decreasing, or a trend of decreasing and increasing, it can be understood that the change characteristics of the loss function are unstable, and the base station determines that it is difficult for the UE to train the first model, and the UE needs to learn an easier first dataset to continue training the first model. If the base station determines that the decrease in the value of the loss function is in line with expectations based on the changing characteristics of the loss function, that is, the value of the loss function gradually decreases and decreases rapidly, the base station determines that the difficulty of the second dataset used to train the first model is appropriate, and the base station can obtain a first dataset with the same difficulty as the second dataset based on the course learning algorithm.

[0188] For example, if the first model is used by the UE to send CSI to the base station, both the first dataset and the second dataset can be understood as CSI matrices. When the difficulty of the first dataset is greater than that of the second dataset, it can be understood that the data distribution of the CSI matrix corresponding to the second dataset is wider, making it easier for the UE to compress the CSI matrix. Consequently, when the base station receives the compressed CSI vector from the UE, it is easier to decompress the CSI vector. Conversely, when the data distribution of the CSI matrix corresponding to the first dataset is narrower, the UE faces greater difficulty in compressing the CSI matrix, making it more difficult for the base station to decompress the compressed CSI vector from the UE. The value of the loss function can be understood as the difference between the CSI vector (predicted value) obtained by the UE compressing the CSI matrix using the first model and the original compressed CSI vector (true value).

[0189] If the second dataset is used as part of the training set, when the UE performs multiple training rounds based on the second dataset, the parameters of the first model can be obtained after each training batch in each training round, and the parameters of the first model can be updated. The next training batch continues training based on the updated first model. When the base station receives one or more performance metrics reported by the UE, it can first determine whether the target performance has been achieved based on these results. For example, the base station can determine whether the first model has achieved the target performance based on the curve of the received loss function. If it is determined that the target performance has been achieved, the base station determines that the first model training is complete, and the base station does not need to determine the first dataset; that is, the base station stops sending datasets to the UE. If it is determined that the target performance has not been achieved, the base station determines the first dataset to be sent in the next batch based on the course learning algorithm.

[0190] If the second dataset is a validation set, when the UE performs multiple training rounds based on the second dataset, it can input the original data from the second dataset into the first model at each preset training round interval to obtain training results. Based on the training results and the original data, it can calculate the cosine similarity or NMSE, but without updating the parameters of the first model. When the base station receives the cosine similarity or NMSE curves obtained from multiple training round intervals, the base station can first determine whether the first model has reached the expected target based on the cosine similarity or NMSE curves. If it is determined that the target performance has been reached, the base station determines that the first model training is complete, and the base station does not need to determine the first dataset; that is, the base station stops sending datasets to the UE. If it is determined that the target performance has not yet been reached, the base station determines the first dataset to be sent in the next batch based on the course learning algorithm.

[0191] For example, the expected target can be understood as a target threshold. The base station can determine the average value of the curves of the loss function, NMSE, or cosine similarity, and compare the average value with the target threshold. If the average value is less than or equal to the target threshold, the base station determines that the target performance has been achieved. Alternatively, the base station can take the last obtained value of the loss function, NMSE, or cosine similarity from the curves and compare it with the target threshold. If the last obtained value of the loss function, NMSE, or cosine similarity is less than or equal to the target threshold, the base station determines that the target performance has been achieved.

[0192] In some embodiments, based on method two or method three, when the amount of data required for subsequent stages in the training of the first model is reduced, or when the increase in the training completion rate of the multiple stages in the training of the first model is increased, the amount of data in the first dataset is less than the amount of data in the second dataset.

[0193] When the amount of data required for each subsequent stage in the training of the first model increases, or when the increase in the training completion rate of each stage in the training of the first model decreases, the amount of data in the first dataset is greater than or equal to the amount of data in the second dataset.

[0194] The amount of data required for each subsequent stage in the training of the first model, or the training completion level of each stage, can be determined by the user experience (UE) based on a loss function, non-maximum mean squared error (NMSE), or cosine similarity. For example, the UE can determine the convergence level of the first model based on the loss function, NMSE, or cosine similarity, and then determine the amount of data required for each subsequent stage or the training completion level based on the convergence level of the first model. Each stage here can be understood as a training epoch interval.

[0195] For example, when the user determines from the loss function curve that the value of the loss function meets expectations, and that the amount of data required for subsequent stages in the training of the first model decreases while the increase in training completion increases, the first model tends to converge, and the amount of data in the first dataset to be sent in the next batch is less than the amount of data in the second dataset. Conversely, when the user determines from the loss function curve that the value of the loss function decreases slowly or the trend is unstable, and that the amount of data required for subsequent stages in the training of the first model increases while the increase in training completion decreases, the first model has not converged, and the amount of data in the first dataset to be sent in the next batch is greater than or equal to the amount of data in the second dataset.

[0196] 607. The base station sends the first dataset to the UE.

[0197] Accordingly, the UE receives the first dataset sent by the base station. The UE then continues training the first model based on the first dataset.

[0198] Steps 603 to 607 are equivalent to completing one iteration of the model.

[0199] Similar to step 603, when the base station sends the first dataset to the UE, it may also send the identifier of the training set and / or the identifier of the first dataset; or, send the identifier of the first model and / or the identifier of the first dataset; or, send the functional identifier of the first model and / or the identifier of the first dataset.

[0200] If the base station determines that the first model training is complete based on step 606, the base station does not need to determine the first dataset in step 606. Step 607 can be replaced by: the base station sending second indication information, which indicates that the first model training is complete or training is paused. Accordingly, the UE receives the second indication information.

[0201] Therefore, in this application, the base station can distribute the dataset in batches. If the base station determines that the model on the UE side has reached the target performance during training, the base station can terminate the model training on the UE side in advance, saving the overhead of transmitting the remaining dataset to the UE. If the base station determines that the UE side needs to continue model training, the base station can adjust the difficulty and / or data volume of the dataset used for model training on the UE side based on the course learning algorithm and the target parameters reported by the UE, so as to accelerate the convergence of the model training on the UE side and improve the efficiency of model training on the UE side.

[0202] In this application, if applied to an O-RAN architecture, the UE-side model can be deployed either on a chip inside the UE or on a device outside the UE. For example, the device outside the UE could be on a host in an over-the-top (OTT) system or in a cloud server.

[0203] The deployment of base station-side models can be implemented on chips inside the base station or in devices outside the base station. For example, devices outside the base station include network-side AI model deployment equipment collectively referred to as intelligent network elements, such as near-real-time RICs. Near-real-time RICs are, for example, located in RAN nodes, such as in CUs or DUs.

[0204] For example, in a dual-end model, the model on the UE side is an AI CSI encoder, and the model on the base station side is an AI CSI decoder.

[0205] For example, the communication system of this application includes intelligent network elements, base stations, UEs, and OTTs. When the method flow of Figure 6 above is applied to this communication system, Figure 7 is a schematic flowchart of a model training method applied in an O-RAN architecture, including the following flow.

[0206] 701. The base station and UE complete the pairing of the dual-end model.

[0207] 702. The intelligent network element sends a third indication message to the base station, and the base station sends a third indication message to the UE. The third indication message is used to indicate the target parameters to be reported during the training of the first model.

[0208] 703. The intelligent network element sends a second dataset to the base station, the base station sends a second dataset to the UE, and the UE sends a second dataset to the OTT. The second dataset is used for training the first model.

[0209] 704. OTT determines the first indication information based on the second dataset. The first indication information is used to indicate the target parameters during the training process of the first model.

[0210] 705. OTT sends the first instruction information to the intelligent network element through the UE and the base station.

[0211] 706. The intelligent network element determines the target parameters based on the first indication information, and determines the first dataset based on the target parameters.

[0212] 707. The intelligent network element sends the first dataset to the OTT through the base station and UE.

[0213] 708. OTT trains the first model based on the first dataset.

[0214] In other words, the dataset obtained by the base station can originate from a different physical entity than the base station itself; for example, the dataset on the base station side can originate from intelligent network elements. Similarly, the training and deployment of AI models by the UE can be performed in a different physical entity than the UE itself; for example, the training and deployment of AI models by the UE can be performed within an OTT (Over-The-Top) application.

[0215] In this O-RAN architecture, the UE (OTT) can also report target parameters during model training, so that the base station (intelligent network element) can distribute the dataset in batches based on the course learning algorithm to guide the UE in model training, which helps to improve the efficiency of model training on the UE.

[0216] It is understood that, in order to achieve the functions in the above embodiments, the network device and terminal device include hardware structures and / or software modules corresponding to perform each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0217] Figures 8 and 9 are schematic diagrams of possible communication devices provided in embodiments of this application. These communication devices can be used to implement the functions of terminals or base stations in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments. In the embodiments of this application, the communication device can be the terminal 120 or 130 shown in Figure 1, or the base station 110 shown in Figure 1, or a module (such as a chip) applied to a terminal or base station.

[0218] As shown in Figure 8, the communication device 800 includes a processing unit 8110 and a transceiver unit 8120. The communication device 800 is used to implement the functions of the terminal device (UE) or network device (base station) in the method embodiments shown in Figures 5 and 6 above.

[0219] When the communication device 800 is used to implement the function of the terminal device in the method embodiment shown in FIG5: the transceiver unit 8120 is used to send first indication information, the first indication information is used to indicate the target parameters in the first model training process; the processing unit 8110 is used to determine the target parameters based on the second dataset.

[0220] When the communication device 800 is used to implement the function of the network device in the method embodiment shown in FIG5: the transceiver unit 8120 is used to receive the first indication information, the first indication information is used to indicate the target parameters in the first model training process; the processing unit 8110 is used to determine the target parameters based on the first indication information.

[0221] When the communication device 800 is used to implement the functions of the UE in the method embodiment shown in FIG. 6: the transceiver unit 8120 is used to receive third indication information, which is used to indicate the reporting of target parameters during the training of the first model; receive a second dataset; send first indication information to the base station; and receive the first dataset. The processing unit 8110 is used to complete the pairing of the dual-end models with the base station; and determine the first indication information based on the second dataset.

[0222] When the communication device 800 is used to implement the functions of the base station in the method embodiment shown in FIG6: the transceiver unit 8120 is used to send third indication information; send a second dataset; receive first indication information; and send a first dataset. The processing unit 8110 is used to complete the pairing of the dual-end model with the UE; determine the target parameters based on the first indication information, and determine the first dataset based on the target parameters.

[0223] For a more detailed description of the above-mentioned processing unit 8110 and transceiver unit 8120, please refer to the relevant descriptions in the method embodiments shown in Figures 5 and 6.

[0224] As shown in Figure 9, the communication device 900 includes a processor 9210 and an interface circuit 9220. The processor 9210 and the interface circuit 9220 are coupled to each other. It is understood that the interface circuit 9220 can be a transceiver or an input / output interface. Optionally, the communication device 900 may also include a memory 9230 for storing instructions executed by the processor 9210, or storing input data required by the processor 9210 to execute instructions, or storing data generated after the processor 9210 executes instructions.

[0225] When the communication device 900 is used to implement the method shown in FIG5 or FIG6, the processor 9210 is used to implement the function of the processing unit 8110, and the interface circuit 9220 is used to implement the function of the transceiver unit 8120.

[0226] When the communication device 900 is a chip within a terminal device, the deployment of various functions on the terminal device side can be within different parts of the chip. Figure 10 shows a schematic diagram of the structure of a chip 1000 within a terminal device provided in this application, including an antenna, a receiver and / or a transmitter, a controller, a memory, and a processor.

[0227] For example, the terminal device establishes a communication connection with the network device through the controller, and the terminal device receives the signaling, reference signals and datasets (such as the first dataset and / or the second dataset in this application) sent by the network device through the antenna and receiver. The terminal device performs model training through the processor, stores the training results of the model in the memory, and sends the target parameters during the training process to the network device through the transmitter and antenna.

[0228] When the aforementioned communication device is a chip applied to a terminal, the terminal chip implements the functions of the terminal in the above method embodiments. The terminal chip receives information from the base station, which can be understood as the information being first received by other modules in the terminal (such as an RF module or antenna), and then sent to the terminal chip by these modules. The terminal chip sends information to the base station, which can be understood as the information being first sent to other modules in the terminal (such as an RF module or antenna), and then sent to the base station by these modules.

[0229] When the aforementioned communication device is a chip applied to a base station, the base station chip implements the functions of the base station in the above method embodiments. The base station chip receives information from the terminal, which can be understood as the information being first received by other modules in the base station (such as an RF module or antenna), and then sent to the base station chip by these modules. The base station chip sends information to the terminal, which can be understood as the information being sent down to other modules in the base station (such as an RF module or antenna), and then sent to the terminal by these modules.

[0230] In this application, entity A sends information to entity B, either directly or indirectly through other entities. Similarly, entity B receives information from entity A, either directly or indirectly through other entities. Entities A and B can be RAN nodes or terminals, or modules within RAN nodes or terminals. Information transmission and reception can be between RAN nodes and terminals, such as between a base station and a terminal; between two RAN nodes, such as between a CU and a DU; or between different modules within a single device, such as between a terminal chip and other modules of the terminal, or between a base station chip and other modules of the base station.

[0231] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0232] The method steps in the embodiments of this application can be implemented in hardware or in software instructions executable by a processor. The software instructions can consist of corresponding software modules, which can be stored in random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. The storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a base station or terminal. The processor and storage medium can also exist as discrete components in a base station or terminal.

[0233] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video optical disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both types of storage media.

[0234] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0235] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. In the textual description of this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. "Including at least one of A, B, and C" can mean: including A; including B; including C; including A and B; including A and C; including B and C; including A, B, and C.

[0236] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The order of the process numbers described above does not imply the order of execution; the execution order of each process should be determined by its function and internal logic.

Claims

1. A model training method, characterized in that, The method includes: Receive first instruction information, the first instruction information being used to indicate target parameters during the training process of the first model, the target parameters including at least one of the following: one or more results of performance metrics during the training process of the first model, the amount of data required for subsequent stages of one or more stages during the training process of the first model, or the training completion progress of one or more stages during the training process of the first model. The target parameters are determined based on the first indication information.

2. The method according to claim 1, characterized in that, The method further includes: Send a first dataset, which is determined based on the target parameters, and use the first dataset for training the first model.

3. The method according to claim 2, characterized in that, The first dataset, determined based on the target parameters, includes: The amount of data and / or the difficulty of the data in the first dataset are determined based on the target parameters.

4. The method according to claim 1, characterized in that, The method further includes: Send a second instruction message, which indicates that the training of the first model is complete or training is paused.

5. The method according to any one of claims 1-4, characterized in that, Before receiving the first indication information, the method further includes: Send a second dataset, the first indication information being determined based on the second dataset, which is used for training the first model.

6. The method according to any one of claims 1-5, characterized in that, The performance metrics include the results of one or more target metrics obtained during a first training phase using the second dataset, wherein the first training phase includes one or more training rounds. The result of the target metric is used to measure the difference between the predicted value and the true value of the first model.

7. The method according to claim 6, characterized in that, The target metric includes at least one of the following: the value of the loss function of the second dataset, the normalized mean square error (NMSE), or the cosine similarity.

8. The method according to any one of claims 5-7, characterized in that, The data difficulty of the first dataset is different from that of the second dataset.

9. The method according to claim 8, characterized in that, When the changes in multiple results of the performance metrics gradually decrease, the data difficulty of the first dataset is greater than that of the second dataset.

10. The method according to any one of claims 5-9, characterized in that, The second dataset serves as the validation set for the first model.

11. The method according to any one of claims 1-5, characterized in that, When the amount of data required for each subsequent stage in the process of training the first model decreases, or when the amount of increase in the training completion rate of each stage in the process of training the first model increases, the amount of data in the first dataset is less than the amount of data in the second dataset. When the amount of data required for each subsequent stage in the training of the first model increases, or when the increase in the training completion rate of each stage in the training of the first model decreases, the amount of data in the first dataset is greater than or equal to the amount of data in the second dataset.

12. The method according to any one of claims 1-11, characterized in that, Before receiving the first indication information, the method further includes: Send a third instruction message, which is used to instruct the target parameters to be reported during the training of the first model.

13. The method according to any one of claims 2-12, characterized in that, The method further includes: Send the identifier of the training set and / or the identifier of the first dataset; Alternatively, send the identifier of the first model and / or the identifier of the first dataset; Alternatively, send the function identifier corresponding to the first model and / or the identifier of the first dataset.

14. A model training method, characterized in that, The method includes: Determine first indication information, which is used to indicate target parameters in the training process of the first model. The target parameters include at least one of the following: one or more results of performance indicators in the training process of the first model, the amount of data required for the subsequent stages of one or more stages in the training process of the first model, or the training completion progress of one or more stages in the training process of the first model. Send the first instruction message.

15. The method according to claim 14, characterized in that, The method further includes: Receive a first dataset, which is determined based on the target parameters, and use the first dataset for training the first model.

16. The method according to claim 15, characterized in that, The first dataset, determined based on the target parameters, includes: The amount of data and / or the difficulty of the data in the first dataset are determined based on the target parameters.

17. The method according to claim 14, characterized in that, The method further includes: Receive a second instruction message, which indicates that the training of the first model is complete or training is paused.

18. The method according to any one of claims 14-17, characterized in that, Before sending the first indication information, the method further includes: receiving a second dataset, the second dataset being used for training the first model; The determination of the first indication information includes: determining the first indication information based on the second dataset.

19. The method according to any one of claims 14-18, characterized in that, The performance metrics include the results of one or more target metrics obtained during a first training phase using the second dataset, wherein the first training phase includes one or more training rounds. The result of the target metric is used to measure the difference between the predicted value and the true value of the first model.

20. The method according to claim 19, characterized in that, The target metric includes at least one of the following: the value of the loss function of the second dataset, the normalized mean square error (NMSE), or the cosine similarity.

21. The method according to any one of claims 18-20, characterized in that, The data difficulty of the first dataset is different from that of the second dataset.

22. The method according to claim 21, characterized in that, When the changes in multiple results of the performance metrics gradually decrease, the data difficulty of the first dataset is greater than that of the second dataset.

23. The method according to any one of claims 18-22, characterized in that, The second dataset serves as the validation set for the first model.

24. The method according to any one of claims 14-18, characterized in that, When the amount of data required for each subsequent stage in the process of training the first model decreases, or when the amount of increase in the training completion rate of each stage in the process of training the first model increases, the amount of data in the first dataset is less than the amount of data in the second dataset. When the amount of data required for each subsequent stage in the training of the first model increases, or when the increase in the training completion rate of each stage in the training of the first model decreases, the amount of data in the first dataset is greater than or equal to the amount of data in the second dataset.

25. The method according to any one of claims 14-24, characterized in that, Before determining the first indication information, the method further includes: Receive third instruction information, which is used to instruct the reporting of the target parameters during the training of the first model.

26. The method according to any one of claims 15-25, characterized in that, The method further includes: Receive the identifier of the training set and / or the identifier of the first dataset; Alternatively, receive the identifier of the first model and / or the identifier of the first dataset; Alternatively, receive the function identifier corresponding to the first model and / or the identifier of the first dataset.

27. A communication device, characterized in that, The communication device includes a module for performing the method as described in any one of claims 1-26.

28. A communication device, characterized in that, The device includes a processor and an interface circuit, wherein the interface circuit is used to receive signals from other communication devices and transmit them to the processor or to send signals from the processor to other communication devices, and the processor is used to implement the method as described in any one of claims 1 to 26 through logic circuits or executing code instructions.

29. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a communication device, cause the communication device to perform the method as described in any one of claims 1-26.

30. A computer program product, characterized in that, Includes computer instructions that, when executed on a communication device, cause the communication device to perform the method as described in any one of claims 1-26.

Citation Information

Patent Citations

  • Communication method and device

    CN115802370A

  • Named entity recognition model training method and device, equipment and medium

    CN116362251A

  • Model training method and device based on communication network, equipment and medium

    CN118820025A

  • System and method for DNN-based cyber-security using federated learning-based generative adversarial network

    US20230308465A1