Data collection method and apparatus, first device, second device

By selecting devices that need to collect training data within the wireless communication system, the problems of transmission pressure and duplicate samples in centralized training are solved, achieving more efficient data collection and model training.

CN116264712BActive Publication Date: 2026-05-12VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2021-12-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In wireless communication systems, centralized training requires the collection of a large amount of data, but the data at remote locations overlaps, leading to increased transmission pressure and too many duplicate samples in the dataset, which affects the efficiency of model training.

Method used

The first device screens candidate second devices to determine which devices need to collect and report training data. Instructions are sent to reduce data transmission pressure and avoid duplicate samples. The second device is instructed to collect and report training data using unicast or broadcast methods.

Benefits of technology

It alleviated the transmission pressure during data collection, avoided excessive duplicate samples in the dataset, reduced the pressure on model training, and improved data diversity and training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116264712B_ABST
    Figure CN116264712B_ABST
Patent Text Reader

Abstract

The application discloses a data collection method and device, a first device, and a second device, and belongs to the technical field of communication. The data collection method of the application embodiment comprises the following steps: a first device sends a first instruction to a second device, and the instruction indicates that the second device collects and reports training data used for specific AI model training; the first device receives the training data reported by the second device; and the first device constructs a data set by using the training data, and trains the specific AI model. The application embodiment can relieve the transmission pressure during data collection, and avoid too many repeated samples in the data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, specifically relating to a data collection method and apparatus, a first device, and a second device. Background Technology

[0002] Artificial intelligence (AI) has been widely applied in various fields. AI modules can be implemented in various ways, such as neural networks, decision trees, support vector machines, and Bayesian classifiers.

[0003] When training AI modules in wireless communication systems in a centralized manner, a large amount of data needs to be collected from remote locations to construct the dataset. Since there is significant overlap in the remote wireless data, it is unnecessary to collect data from all users. This alleviates the transmission pressure during data collection and avoids an excessive number of duplicate (or similar) samples in the dataset. Summary of the Invention

[0004] This application provides a data collection method and apparatus, a first device, and a second device, which can alleviate the transmission pressure during data collection and avoid too many duplicate samples in the data.

[0005] Firstly, a data collection method is provided, including:

[0006] The first device sends a first instruction to the second device, instructing the second device to collect and report training data used for training a specific AI model;

[0007] The first device receives the training data reported by the second device;

[0008] The first device uses the training data to construct a dataset and train the specific AI model.

[0009] Secondly, a data collection device is provided, comprising:

[0010] The sending module is used to send a first instruction to the second device, instructing the second device to collect and report training data for training a specific AI model;

[0011] A receiving module is used to receive training data reported by the second device;

[0012] The training module is used to construct a dataset using the training data and train the specific AI model.

[0013] Thirdly, a data collection method is provided, including:

[0014] The second device receives a first instruction from the first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model;

[0015] The second device collects training data and reports the training data to the first device.

[0016] Fourthly, a data collection device is provided, comprising:

[0017] The receiving module is configured to receive a first instruction from the first device, wherein the first instruction is used to instruct the second device to collect and report training data for training a specific AI model.

[0018] The processing module is used to collect training data and report the training data to the first device.

[0019] Fifthly, a first device is provided, the first device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.

[0020] In a sixth aspect, a first device is provided, including a processor and a communication interface, wherein the communication interface is used to send a first instruction to a second device, instructing the second device to collect and report training data for training a specific AI model; and to receive the training data reported by the second device; the processor is used to construct a dataset using the training data and to train the specific AI model.

[0021] In a seventh aspect, a second device is provided, the second device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the third aspect.

[0022] Eighthly, a second device is provided, including a processor and a communication interface, wherein the communication interface is used to receive a first instruction from a first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model; the processor is used to collect the training data and report the training data to the first device.

[0023] A ninth aspect provides a data collection system, comprising: a first device and a second device, wherein the first device is configured to perform the steps of the data collection method as described in the first aspect, and the second device is configured to perform the steps of the data collection method as described in the third aspect.

[0024] In a tenth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the third aspect.

[0025] Eleventhly, a chip is provided, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect, or to implement the method as described in the third aspect.

[0026] In a twelfth aspect, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the data collection method as described in the first aspect, or to implement the steps of the data collection method as described in the third aspect.

[0027] In this embodiment, the first device does not require all candidate second devices to collect and report training data. Instead, the first device first filters the candidate second devices, determines the second devices that need to collect and report training data, and then sends a first instruction to the second device to collect and report training data. This can alleviate the transmission pressure during data collection and avoid too many duplicate samples in the dataset, thus reducing the pressure on model training. Attached Figure Description

[0028] Figure 1 This is a block diagram of a wireless communication system applicable to embodiments of this application;

[0029] Figure 2 This is a schematic diagram of channel state information feedback;

[0030] Figure 3 This is a performance diagram illustrating the AI ​​training process at different iteration numbers;

[0031] Figure 4 This is a schematic flowchart of the data collection method on the device side in the first embodiment of this application;

[0032] Figure 5 This is a schematic flowchart of the second device-side data collection method according to an embodiment of this application;

[0033] Figure 6 This is a schematic diagram of the structure of the communication device according to an embodiment of this application;

[0034] Figure 7 This is a schematic diagram of the terminal structure according to an embodiment of this application;

[0035] Figure 8 This is a schematic diagram of the network-side device in an embodiment of this application. Detailed Implementation

[0036] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0037] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0038] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency Division Multiple Access (SC-FDMA), and other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used with the systems and radio technologies mentioned above, as well as with other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and NR terminology is used in most of the following description; however, these technologies can also be applied to applications beyond NR systems, such as 6th generation (6G) radio systems. th Generation 6G communication system.

[0039] Figure 1This diagram illustrates a block diagram of a wireless communication system applicable to embodiments of this application. The wireless communication system includes a terminal 11 and a network-side device 12. Terminal 11 can be a mobile phone, tablet computer, laptop computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, vehicle-mounted device (VUE), pedestrian terminal (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. It should be noted that the specific type of terminal 11 is not limited in this embodiment. Network-side equipment 12 may include access network equipment or core network equipment. Access network equipment 12 may also be referred to as radio access network equipment, radio access network (RAN), radio access network function, or radio access network unit. Access network equipment 12 may include base stations, WLAN access points, or WiFi nodes, etc. Base stations may be referred to as Node B, evolved Node B (eNB), access point, base transceiver station (BTS), radio base station, radio transceiver, Basic Service Set (BSS), Extended Service Set (ESS), home B node, home evolved B node, Transmitting Receiving Point (TRP), or any other suitable term in the field, as long as the same technical effect is achieved. The base station is not limited to specific technical terms. It should be noted that in this application embodiment, only a base station in an NR system is used as an example for description, and the specific type of base station is not limited.

[0040] Generally, the AI ​​algorithms and models selected vary depending on the type of problem. The main method for improving 5G network performance using AI is to enhance or replace existing algorithms or processing modules with neural network-based algorithms and models. In specific scenarios, neural network-based algorithms and models can achieve better performance than deterministic algorithms. Commonly used neural networks include deep neural networks, convolutional neural networks, and recurrent neural networks. Existing AI tools can be used to build, train, and validate neural networks.

[0041] Replacing modules in existing systems with AI methods can effectively improve system performance. For example... Figure 2 The Channel State Information (CSI) feedback shown can be significantly improved in terms of system performance by replacing conventional CSI calculations with AI encoders and decoders, while maintaining the same overhead. Through this AI-based approach, the system's spectral efficiency can be improved by approximately 30%.

[0042] Performance of AI training at different iterations, such as Figure 3 As shown in the figure, the horizontal axis represents the training period, and the vertical axis represents the square of the correlation. Different iterations require different training data, and it can be seen that a large number of training iterations are needed to achieve performance convergence.

[0043] When AI is applied to wireless communication systems, many models operate on the terminal side, and many tasks on the base station side require training using data collected from the terminal side. Due to the limited computing power of the terminal itself, a feasible solution is to report the data to the network side for centralized training. Since many terminals share similar terminal types and service types, and operate in similar environments, their data tends to have high similarity. The training dataset should avoid generating identical data as much as possible, as identical data offers little benefit to model convergence and may even lead to overfitting or poor generalization performance. Meta-learning is a learning method to improve model generalization ability. It constructs multiple tasks based on a dataset, performs multi-task learning, and obtains an optimal initial model. In new scenarios, the initial model obtained through meta-learning can quickly achieve fine-tuning and convergence, exhibiting high adaptability. Constructing multiple tasks requires a certain degree of difference in data features between different tasks. Therefore, both traditional learning schemes and meta-learning schemes require a certain degree of overlap or similarity in the data within the dataset.

[0044] The data collection method provided in this application will be described in detail below with reference to the accompanying drawings, through some embodiments and application scenarios.

[0045] This application provides a data collection method, such as... Figure 4 As shown, it includes:

[0046] Step 101: The first device sends a first instruction to the second device, instructing the second device to collect and report training data used for training a specific AI model;

[0047] Step 102: The first device receives the training data reported by the second device;

[0048] Step 103: The first device uses the training data to construct a dataset and train the specific AI model.

[0049] In this embodiment, the first device does not require all candidate second devices to collect and report training data. Instead, the first device first filters the candidate second devices, determines the second devices that need to collect and report training data, and then sends a first instruction to the second device to collect and report training data. This can alleviate the transmission pressure during data collection and avoid too many duplicate samples in the dataset, thus reducing the pressure on model training.

[0050] In some embodiments, the first device sending a first instruction to the second device includes:

[0051] The first device selects N second devices from M candidate second devices according to a preset first filtering condition, and unicasts the first instruction to the N second devices, where M and N are positive integers, and N is less than or equal to M; or

[0052] The first device broadcasts the first instruction to the M candidate second devices. The first instruction carries a second filtering condition, which is used to filter the second devices that report the training data. The second devices meet the second filtering condition.

[0053] In this embodiment, devices within the communication range of the first device are considered candidate second devices. The second device reporting training data is selected from these candidate second devices; all candidate second devices can be used, or a subset can be selected. Broadcasting sends a first instruction to all candidate second devices, while unicasting only sends the first instruction to the selected second devices. Candidate second devices receiving the unicast first instruction are required to collect and report training data. Candidate second devices receiving the broadcast first instruction need to determine whether they meet the second selection criteria; only those meeting the second selection criteria collect and report training data.

[0054] In some embodiments, the first device sends a first instruction to the second device via at least one of the following:

[0055] Media intervention control MAC control unit CE;

[0056] Radio Resource Control (RRC) messages;

[0057] Non-access stratum NAS messages;

[0058] Manage and orchestrate messages;

[0059] User face data;

[0060] Downlink control information;

[0061] System Information Block (SIB);

[0062] Layer 1 signaling of the Physical Downlink Control Channel (PDCCH);

[0063] Information about the Physical Downlink Shared Channel (PDSCH);

[0064] MSG 2 information of the Physical Random Access Channel (PRACH);

[0065] MSG 4 information of the Physical Random Access Channel (PRACH);

[0066] MSG B information of the Physical Random Access Channel (PRACH);

[0067] Information or signaling broadcast from the channel;

[0068] Xn interface signaling;

[0069] PC5 interface signaling;

[0070] Information or signaling from the Physical Side Link Control Channel (PSCCH);

[0071] Information about the Physical Side Link Shared Channel (PSSCH);

[0072] Information from the Physical Side Link Broadcast Channel (PSBCH);

[0073] Information about the Physical Straight-Through Link Discovery Channel (PSDCH);

[0074] Information from the Physical Straight-Through Link Feedback Channel (PSFCH).

[0075] In some embodiments, before the first device sends a first instruction to the second device, the method further includes:

[0076] The first device receives the first training data and / or the first parameter reported by the candidate second device, wherein the first parameter may be the judgment parameter of the first screening condition.

[0077] In this embodiment, the candidate second device can first report a small amount of training data (i.e., the first training data) and / or the first parameters. The first device determines the data source for training based on the small amount of training data and / or the first parameters, and selects the second device that collects and reports training data, thus avoiding all second devices reporting training data.

[0078] In some embodiments, the first device only receives the first training data reported by the candidate second devices, and determines the first parameter based on the first training data. The first device can infer, sense, detect, or deduce the first parameter based on the first training data. The first device can then screen candidate second devices based on the first parameter to determine the second device.

[0079] In some embodiments, the first parameter includes at least one of the following:

[0080] The data type of the candidate second device;

[0081] The data distribution parameters of the candidate second device;

[0082] The service types of the candidate second devices include, for example, enhanced mobile broadband (eMBB), ultra-reliable low-latency communication (URLLC), massive machine-type communication (mMTC), and other new 6G scenarios.

[0083] The operating scenarios of the candidate second device include, but are not limited to: high speed, low speed, line-of-sight (LOS) propagation, non-line-of-sight (NLOS) propagation, high signal-to-noise ratio (SNR), and low signal-to-noise ratio (SNR) operating scenarios.

[0084] The communication network access methods of the candidate second device include mobile network, WiFi and fixed network, wherein the mobile network includes 2G, 3G, 4G, 5G and 6G;

[0085] The channel quality of the candidate second device;

[0086] The degree of difficulty in collecting data by the candidate second device;

[0087] The power status of the candidate second device, such as the specific value of the remaining available power, or the hierarchical description result, such as charging or not charging;

[0088] The storage status of the candidate second device, such as the specific value of available memory, or the hierarchical description result.

[0089] In this embodiment, the candidate second device may first report a small amount of training data (i.e., the first training data) and / or the first parameter to the first device. The first parameter may be a judgment parameter of the first screening condition. The first device determines the second device that needs to collect and report training data according to the first screening condition. The second device is selected from the candidate second devices. Specifically, there may be M candidate second devices, from which N second devices are determined that need to collect and report training data. N may be less than M or equal to M.

[0090] In a specific example, the second devices required for training data collection and reporting can be determined based on their data types. The candidate second devices are then grouped according to their data types, with each group containing devices of the same or similar data types. When selecting the second devices, K1 candidate second devices are chosen from each group as the devices required for training data collection and reporting, where K1 is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0091] In a specific example, the second devices that need to collect and report training data can be determined based on their service types. The candidate second devices are then grouped according to their service types, with each group containing devices of the same or similar service types. When selecting the second devices, K2 candidate second devices are chosen from each group as the devices required to collect and report training data, where K2 is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0092] In a specific example, the second devices that need to collect and report training data can be determined based on the data distribution parameters of the candidate second devices. The candidate second devices are then grouped according to their data distribution parameters, with each group containing devices of similar or identical data distribution parameters. When selecting the second devices, K3 candidate second devices are chosen from each group as the second devices required to collect and report training data, where K3 is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0093] In a specific example, the second devices required for training data collection and reporting can be determined based on their operating scenarios. The candidate second devices are then grouped according to their operating scenarios, with each group containing devices in the same or similar scenarios. When selecting the second devices, A candidate second devices are chosen from each group as the devices required for training data collection and reporting, where A is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0094] In a specific example, the second devices required for training data collection and reporting can be determined based on their communication network access methods. These second devices are then prioritized according to their communication network access methods, which include fixed-line, WiFi, and mobile networks. Mobile networks include 2G, 3G, 4G, 5G, and 6G. Fixed-line devices have a higher priority than WiFi devices, and WiFi devices have a higher priority than mobile network devices. Within mobile networks, higher algebraic numbers result in higher priority; for example, 5G candidates have a higher priority than 4G candidates. B candidate devices, where B is a positive integer, are selected from these priority groups in descending order as the second devices required for training data collection and reporting.

[0095] In a specific example, the second devices that need to collect and report training data can be determined based on the channel quality of the candidate second devices. The candidate second devices are then prioritized according to their channel quality, with higher-quality candidate second devices having higher priority. C candidate second devices are selected from the candidates in descending order of priority as the second devices that need to collect and report training data, where C is a positive integer. This ensures that the second devices with good channel quality collect and report training data, thus guaranteeing the training quality of the specific AI model.

[0096] In a specific example, the second device that needs to collect and report training data can be determined based on the difficulty of collecting data from the candidate second devices. The candidate second devices are prioritized according to the difficulty of collecting data. The candidate second device with the easier data collection is selected, the higher its priority. D candidate second devices are selected from the candidate second devices in descending order of priority as the second devices that need to collect and report training data, where D is a positive integer. This can reduce the difficulty of data collection.

[0097] In a specific example, the second devices that need to collect and report training data can be determined based on the battery status of the candidate second devices. The candidate second devices are prioritized according to their battery status, with higher battery levels resulting in higher priority. In addition, candidate second devices that are charging have the highest priority. E candidate second devices are selected from the candidates in descending order of priority as the second devices that need to collect and report training data, where E is a positive integer. This ensures that the second devices that collect and report training data have sufficient battery power.

[0098] In a specific example, the second devices that need to collect and report training data can be determined based on the storage status of the candidate second devices. The candidate second devices are prioritized according to their storage status. The larger the available storage space of the candidate second devices, the higher the priority of the selected devices. F candidate second devices are selected from the candidate second devices in descending order of priority as the second devices that need to collect and report training data, where F is a positive integer. This ensures that the second devices that collect and report training data have enough available storage space to store the training data.

[0099] In some embodiments, the first instruction of unicast includes at least one of the following:

[0100] The number of training data samples collected by the second device may be different or the same for different second devices;

[0101] The time it takes for the second device to collect training data may vary or be the same for different second devices.

[0102] The time at which the second device reports training data to the first device may vary or be the same for different second devices.

[0103] Does the collected data need to be preprocessed?

[0104] Methods for preprocessing the collected data;

[0105] The data format of the training data reported by the second device to the first device.

[0106] In some embodiments, the first instruction broadcast includes at least one of the following:

[0107] Identification of the candidate second device for data collection;

[0108] Identifier of a candidate second device that does not collect data;

[0109] The number of training data samples that the candidate second device needs to collect for data collection; the number of training data samples collected by different candidate second devices may be different or the same.

[0110] The time taken for candidate second devices to collect training data; the time taken for different candidate second devices to collect training data may be different or the same.

[0111] The time at which the candidate second device for data collection reports training data to the first device may vary or be the same for different candidate second devices.

[0112] Does the collected data need to be preprocessed?

[0113] Methods for preprocessing the collected data;

[0114] The data format of the training data reported by the candidate second device for data collection to the first device;

[0115] The first filtering condition.

[0116] The identifiers of candidate second devices that collect data and the identifiers of candidate second devices that do not collect data constitute the second screening condition. Candidate second devices can determine whether they meet the second screening condition based on their own identifiers.

[0117] In some embodiments, after training a specific AI model, the method further includes:

[0118] The first device sends the trained AI model and hyperparameters to L inference devices, where L is greater than M, equal to M, or less than M.

[0119] In this embodiment, the first device constructs a training dataset based on the received training data, trains a specific AI model, and sends the converged AI model and hyperparameters to L inference devices. The inference devices are the second devices that need to perform performance verification and inference on the AI ​​model. The inference devices can be selected from the candidate second devices or other second devices besides the candidate second devices.

[0120] In some embodiments, the first device sends the trained AI model and hyperparameters to the inference device via at least one of the following:

[0121] Media intervention control MAC control unit CE;

[0122] Radio Resource Control (RRC) messages;

[0123] Non-access stratum NAS messages;

[0124] Manage and orchestrate messages;

[0125] User face data;

[0126] Downlink control information;

[0127] System Information Block (SIB);

[0128] Layer 1 signaling of the Physical Downlink Control Channel (PDCCH);

[0129] Information about the Physical Downlink Shared Channel (PDSCH);

[0130] MSG 2 information of the Physical Random Access Channel (PRACH);

[0131] MSG 4 information of the Physical Random Access Channel (PRACH);

[0132] MSG B information of the Physical Random Access Channel (PRACH);

[0133] Information or signaling broadcast from the channel;

[0134] Xn interface signaling;

[0135] PC5 interface signaling;

[0136] Information or signaling from the Physical Side Link Control Channel (PSCCH);

[0137] Information about the Physical Side Link Shared Channel (PSSCH);

[0138] Information from the Physical Side Link Broadcast Channel (PSBCH);

[0139] Information about the Physical Straight-Through Link Discovery Channel (PSDCH);

[0140] Information from the Physical Straight-Through Link Feedback Channel (PSFCH).

[0141] In some embodiments, the AI ​​model is a meta-learning model, and the hyperparameters can be determined by the first parameter.

[0142] In some embodiments, the hyperparameters associated with the meta-learning model include at least one of the following:

[0143] External learning rate;

[0144] The internal iterative learning rate corresponding to different training tasks or the inference device;

[0145] Meta-learning rate;

[0146] The number of internal iterations corresponding to different training tasks or the inference device;

[0147] The number of external iterations corresponding to different training tasks or the inference device.

[0148] In this embodiment, the first device can be a network-side device and the second device can be a terminal; or, the first device can be a network-side device and the second device can be a network-side device, such as in a scenario where multiple network-side devices aggregate training data into one network-side device for training; or, the first device can be a terminal and the second device can be a terminal, such as in a scenario where multiple terminals aggregate training data into one terminal for training.

[0149] In addition, the candidate second device can be a network-side device or a terminal; the inference device can be a network-side device or a terminal.

[0150] This application also provides a data collection method, such as... Figure 5 As shown, it includes:

[0151] Step 201: The second device receives a first instruction from the first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model;

[0152] Step 202: The second device collects training data and reports the training data to the first device.

[0153] In this embodiment, the first device does not require all candidate second devices to collect and report training data. Instead, the first device first filters the candidate second devices, determines the second devices that need to collect and report training data, and then sends a first instruction to the second device to collect and report training data. This can alleviate the transmission pressure during data collection and avoid too many duplicate samples in the dataset, thus reducing the pressure on model training.

[0154] In some embodiments, the second device reports training data to the first device by at least one of the following:

[0155] Media intervention control MAC control unit CE;

[0156] Radio Resource Control (RRC) messages;

[0157] Non-access stratum NAS messages;

[0158] Layer 1 signaling of the Physical Uplink Control Channel (PUCCH);

[0159] MSG 1 information of the Physical Random Access Channel (PRACH);

[0160] MSG 3 information of the Physical Random Access Channel (PRACH);

[0161] MSG A information of the Physical Random Access Channel (PRACH);

[0162] Information about the Physical Uplink Shared Channel (PUSCH);

[0163] Xn interface signaling;

[0164] PC5 interface signaling;

[0165] Information or signaling from the Physical Side Link Control Channel (PSCCH);

[0166] Information about the Physical Side Link Shared Channel (PSSCH);

[0167] Information from the Physical Side Link Broadcast Channel (PSBCH);

[0168] Information about the Physical Straight-Through Link Discovery Channel (PSDCH);

[0169] Information from the Physical Straight-Through Link Feedback Channel (PSFCH).

[0170] In some embodiments, the second device receiving a first instruction from the first device includes:

[0171] The second device receives the first instruction unicast by the first device, and the second device is the second device selected by the first device from the candidate second devices according to a preset first filtering condition; or

[0172] The second device receives the first instruction broadcast by the first device to the candidate second device. The first instruction carries a second filtering condition, which is used to filter the second devices that report the training data. The second device meets the second filtering condition.

[0173] In some embodiments, the second device collects training data and reports the training data to the first device, including:

[0174] If the second device receives the first instruction unicast by the first device, the second device collects and reports the training data; or

[0175] If the second device receives the first instruction broadcast by the first device, the second device collects and reports the training data.

[0176] In this embodiment, devices within the communication range of the first device are considered candidate second devices. The second device reporting training data is selected from these candidate second devices; all candidate second devices can be used, or a subset can be selected. Broadcasting sends a first instruction to all candidate second devices, while unicasting only sends the first instruction to the selected second devices. Candidate second devices receiving the unicast first instruction are required to collect and report training data. Candidate second devices receiving the broadcast first instruction need to determine whether they meet the second selection criteria; only those meeting the second selection criteria collect and report training data.

[0177] In some embodiments, before the second device receives a first instruction from the first device, the method further includes:

[0178] The candidate second device reports the first training data and / or the first parameter to the first device, where the first parameter may be a judgment parameter of the first screening condition.

[0179] In this embodiment, the candidate second device can first report a small amount of training data (i.e., the first training data) and / or the first parameters. The first device determines the data source for training based on the small amount of training data and / or the first parameters, and selects the second device that collects and reports training data, thus avoiding all second devices reporting training data.

[0180] In some embodiments, the second device reports first training data and / or first parameters to the first device via at least one of the following:

[0181] Media intervention control MAC control unit CE;

[0182] Radio Resource Control (RRC) messages;

[0183] Non-access stratum NAS messages;

[0184] Layer 1 signaling of the Physical Uplink Control Channel (PUCCH);

[0185] MSG 1 information of the Physical Random Access Channel (PRACH);

[0186] MSG 3 information of the Physical Random Access Channel (PRACH);

[0187] MSG A information of the Physical Random Access Channel (PRACH);

[0188] Information about the Physical Uplink Shared Channel (PUSCH);

[0189] Xn interface signaling;

[0190] PC5 interface signaling;

[0191] Information or signaling from the Physical Side Link Control Channel (PSCCH);

[0192] Information about the Physical Side Link Shared Channel (PSSCH);

[0193] Information from the Physical Side Link Broadcast Channel (PSBCH);

[0194] Information about the Physical Straight-Through Link Discovery Channel (PSDCH);

[0195] Information from the Physical Straight-Through Link Feedback Channel (PSFCH).

[0196] In some embodiments, the candidate second device reports only the first training data to the first device, and the first training data is used to determine the first parameter.

[0197] In some embodiments, the first device only receives the first training data reported by the candidate second devices, and determines the first parameter based on the first training data. The first device can infer, sense, detect, or deduce the first parameter based on the first training data. The first device can then screen candidate second devices based on the first parameter to determine the second device.

[0198] In some embodiments, the first parameter includes at least one of the following:

[0199] The data type of the candidate second device;

[0200] The data distribution parameters of the candidate second device;

[0201] The service types of the candidate second devices include, for example, enhanced mobile broadband (eMBB), ultra-reliable low-latency communication (URLLC), massive machine-type communication (mMTC), and other new 6G scenarios.

[0202] The operating scenarios of the candidate second device include, but are not limited to: high speed, low speed, line-of-sight (LOS) propagation, non-line-of-sight (NLOS) propagation, high signal-to-noise ratio (SNR), and low signal-to-noise ratio (SNR) operating scenarios.

[0203] The communication network access methods of the candidate second device include mobile network, WiFi and fixed network, wherein the mobile network includes 2G, 3G, 4G, 5G and 6G;

[0204] The channel quality of the candidate second device;

[0205] The degree of difficulty in collecting data by the candidate second device;

[0206] The power status of the candidate second device, such as the specific value of the remaining available power, or the hierarchical description result, such as charging or not charging;

[0207] The storage status of the candidate second device, such as the specific value of available memory, or the hierarchical description result.

[0208] In this embodiment, the candidate second device may first report a small amount of training data (i.e., the first training data) and / or the first parameter to the first device. The first parameter may be a judgment parameter of the first screening condition. The first device determines the second device that needs to collect and report training data according to the first screening condition. The second device is selected from the candidate second devices. Specifically, there may be M candidate second devices, from which N second devices are determined that need to collect and report training data. N may be less than M or equal to M.

[0209] In a specific example, the second devices required for training data collection and reporting can be determined based on their data types. The candidate second devices are then grouped according to their data types, with each group containing devices of the same or similar data types. When selecting the second devices, K1 candidate second devices are chosen from each group as the devices required for training data collection and reporting, where K1 is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0210] In a specific example, the second devices that need to collect and report training data can be determined based on their service types. The candidate second devices are then grouped according to their service types, with each group containing devices of the same or similar service types. When selecting the second devices, K2 candidate second devices are chosen from each group as the devices required to collect and report training data, where K2 is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0211] In a specific example, the second devices that need to collect and report training data can be determined based on their data distribution parameters. The candidate second devices are then grouped according to their data distribution parameters, with each group containing devices of similar or identical data distribution parameters. When selecting the second devices, K4 candidate second devices are chosen from each group as those requiring training data collection and reporting, where K3 is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0212] In a specific example, the second devices required for training data collection and reporting can be determined based on their operating scenarios. The candidate second devices are then grouped according to their operating scenarios, with each group containing devices in the same or similar scenarios. When selecting the second devices, A candidate second devices are chosen from each group as the devices required for training data collection and reporting, where A is a positive integer. This ensures the diversity of the training data and guarantees that each group of candidate second devices has a device to collect and report training data.

[0213] In a specific example, the second devices required for training data collection and reporting can be determined based on their communication network access methods. These second devices are then prioritized according to their communication network access methods, which include fixed-line, WiFi, and mobile networks. Mobile networks include 2G, 3G, 4G, 5G, and 6G. Fixed-line devices have a higher priority than WiFi devices, and WiFi devices have a higher priority than mobile network devices. Within mobile networks, higher algebraic numbers result in higher priority; for example, 5G candidates have a higher priority than 4G candidates. B candidate devices, where B is a positive integer, are selected from these priority groups in descending order as the second devices required for training data collection and reporting.

[0214] In a specific example, the second devices that need to collect and report training data can be determined based on the channel quality of the candidate second devices. The candidate second devices are then prioritized according to their channel quality, with higher-quality candidate second devices having higher priority. C candidate second devices are selected from the candidates in descending order of priority as the second devices that need to collect and report training data, where C is a positive integer. This ensures that the second devices with good channel quality collect and report training data, thus guaranteeing the training quality of the specific AI model.

[0215] In a specific example, the second device that needs to collect and report training data can be determined based on the difficulty of collecting data from the candidate second devices. The candidate second devices are prioritized according to the difficulty of collecting data. The candidate second device with the easier data collection is selected, the higher its priority. D candidate second devices are selected from the candidate second devices in descending order of priority as the second devices that need to collect and report training data, where D is a positive integer. This can reduce the difficulty of data collection.

[0216] In a specific example, the second devices that need to collect and report training data can be determined based on the battery status of the candidate second devices. The candidate second devices are prioritized according to their battery status, with higher battery levels resulting in higher priority. In addition, candidate second devices that are charging have the highest priority. E candidate second devices are selected from the candidates in descending order of priority as the second devices that need to collect and report training data, where E is a positive integer. This ensures that the second devices that collect and report training data have sufficient battery power.

[0217] In a specific example, the second devices that need to collect and report training data can be determined based on the storage status of the candidate second devices. The candidate second devices are prioritized according to their storage status. The larger the available storage space of the candidate second devices, the higher the priority of the selected devices. F candidate second devices are selected from the candidate second devices in descending order of priority as the second devices that need to collect and report training data, where F is a positive integer. This ensures that the second devices that collect and report training data have enough available storage space to store the training data.

[0218] In some embodiments, before reporting the training data to the first device, the method further includes:

[0219] The second device sends a first request to the first device, requesting the collection and reporting of training data.

[0220] In some embodiments, the first instruction of unicast includes at least one of the following:

[0221] The number of training data samples collected by the second device may be different or the same for different second devices;

[0222] The time it takes for the second device to collect training data may vary or be the same for different second devices.

[0223] The time at which the second device reports training data to the first device may vary or be the same for different second devices.

[0224] Does the collected data need to be preprocessed?

[0225] Methods for preprocessing the collected data;

[0226] The data format of the training data reported by the second device to the first device.

[0227] In some embodiments, the first instruction broadcast includes at least one of the following:

[0228] Identification of the candidate second device for data collection;

[0229] Identifier of a candidate second device that does not collect data;

[0230] The number of training data samples that the candidate second device needs to collect for data collection; the number of training data samples collected by different candidate second devices may be different or the same.

[0231] The time taken for candidate second devices to collect training data; the time taken for different candidate second devices to collect training data may be different or the same.

[0232] The time at which the candidate second device for data collection reports training data to the first device may vary or be the same for different candidate second devices.

[0233] Does the collected data need to be preprocessed?

[0234] Methods for preprocessing the collected data;

[0235] The data format of the training data reported by the candidate second device for data collection to the first device;

[0236] The first filtering condition.

[0237] The identifiers of candidate second devices that collect data and the identifiers of candidate second devices that do not collect data constitute the second screening condition. Candidate second devices can determine whether they meet the second screening condition based on their own identifiers.

[0238] In some embodiments, after the second device collects training data and reports the training data to the first device, the method further includes:

[0239] The inference device receives the trained AI model and hyperparameters sent by the first device.

[0240] In this embodiment, the first device constructs a training dataset based on the received training data, trains a specific AI model, and sends the converged AI model and hyperparameters to L inference devices. The inference devices are the second devices that need to perform performance verification and inference on the AI ​​model. The inference devices can be selected from the candidate second devices or other second devices besides the candidate second devices.

[0241] In some embodiments, the AI ​​model is a meta-learning model, and the hyperparameters can be determined by the first parameter.

[0242] In some embodiments, the hyperparameters include at least one of the following:

[0243] External learning rate;

[0244] The internal iterative learning rate corresponding to different training tasks or the inference device;

[0245] Meta-learning rate;

[0246] The number of internal iterations corresponding to different training tasks or the inference device;

[0247] The number of external iterations corresponding to different training tasks or the inference device.

[0248] In some embodiments, after the inference device receives the trained AI model and hyperparameters sent by the first device, the method further includes:

[0249] The inference device performs performance verification on the AI ​​model;

[0250] If the performance verification result meets a preset first condition, the inference device uses the AI ​​model for inference. The first condition may be configured, pre-configured, or agreed upon by the first device. After performing performance verification on the AI ​​model, the inference device may also report the result of whether or not inference is performed to the first device.

[0251] In some embodiments, the AI ​​model used for performance verification is the AI ​​model issued by the first device, or the AI ​​model issued by the first device after fine-tuning.

[0252] In this embodiment, the inference device can directly use the AI ​​model issued by the first device for performance verification, or it can fine-tune the AI ​​model issued by the first device before performance verification. For the fine-tuning of meta-learning, the specific hyperparameters related to meta-learning for each inference device can be different. The specific hyperparameters related to meta-learning for each inference device can be determined based on the first parameters corresponding to each inference device (mainly based on the ease of data collection, battery status, storage status, etc. in the first parameters).

[0253] In this embodiment, the first device can be a network-side device and the second device can be a terminal; or, the first device can be a network-side device and the second device can be a network-side device, such as in a scenario where multiple network-side devices aggregate training data into one network-side device for training; or, the first device can be a terminal and the second device can be a terminal, such as in a scenario where multiple terminals aggregate training data into one terminal for training.

[0254] In addition, the candidate second device can be a network-side device or a terminal; the inference device can be a network-side device or a terminal.

[0255] In the above embodiments, the specific AI model can be a channel estimation model, a mobility prediction model, etc. The technical solutions of this application can be applied to 6G networks, as well as 5G and 5.5G networks.

[0256] The data collection method provided in this application can be executed by a data collection device. This application uses an example of a data collection device executing the data collection method to illustrate the data collection device provided in this application.

[0257] This application provides a data collection device, including:

[0258] The sending module is used to send a first instruction to the second device, instructing the second device to collect and report training data for training a specific AI model;

[0259] A receiving module is used to receive training data reported by the second device;

[0260] The training module is used to construct a dataset using the training data and train the specific AI model.

[0261] In some embodiments, the sending module is specifically used to select N second devices from M candidate second devices according to a preset first filtering condition, and unicast the first instruction to the N second devices, where M and N are positive integers, and N is less than or equal to M; or

[0262] The first instruction is broadcast to the M candidate second devices. The first instruction carries a second filtering condition. The second filtering condition is used to filter the second devices that report the training data. The second devices meet the second filtering condition.

[0263] In some embodiments, the receiving module is further configured to receive first training data and / or first parameters reported by the candidate second device, wherein the first parameters may be judgment parameters of the first screening condition.

[0264] In some embodiments, the receiving module is configured to receive only the first training data reported by the candidate second device, and determine the first parameter based on the first training data.

[0265] In some embodiments, the first parameter includes at least one of the following:

[0266] The data type of the candidate second device;

[0267] The data distribution parameters of the candidate second device;

[0268] The service type of the candidate second device;

[0269] The working scenario of the candidate second device;

[0270] The communication network access method of the candidate second device;

[0271] The channel quality of the candidate second device;

[0272] The degree of difficulty in collecting data by the candidate second device;

[0273] The power status of the candidate second device;

[0274] The storage state of the candidate second device.

[0275] In some embodiments, the first instruction of unicast includes at least one of the following:

[0276] The number of training data samples collected by the second device may be different or the same for different second devices;

[0277] The time it takes for the second device to collect training data may vary or be the same for different second devices.

[0278] The time at which the second device reports training data to the first device may vary or be the same for different second devices.

[0279] Does the collected data need to be preprocessed?

[0280] Methods for preprocessing the collected data;

[0281] The data format of the training data reported by the second device to the first device.

[0282] In some embodiments, the first instruction broadcast includes at least one of the following:

[0283] Identification of the candidate second device for data collection;

[0284] Identifier of a candidate second device that does not collect data;

[0285] The number of training data samples that the candidate second device needs to collect for data collection; the number of training data samples collected by different candidate second devices may be different or the same.

[0286] The time taken for candidate second devices to collect training data; the time taken for different candidate second devices to collect training data may be different or the same.

[0287] The time at which the candidate second device for data collection reports training data to the first device may vary or be the same for different candidate second devices.

[0288] Does the collected data need to be preprocessed?

[0289] Methods for preprocessing the collected data;

[0290] The data format of the training data reported by the candidate second device for data collection to the first device;

[0291] The first filtering condition.

[0292] In some embodiments, the sending module is further configured to send the trained AI model and hyperparameters to L inference devices, wherein L is greater than M, equal to M, or less than M.

[0293] In some embodiments, the AI ​​model is a meta-learning model, and the hyperparameters are determined by the first parameter.

[0294] In some embodiments, the hyperparameters include at least one of the following:

[0295] External learning rate;

[0296] The internal iterative learning rate corresponding to different training tasks or the inference device;

[0297] Meta-learning rate;

[0298] The number of internal iterations corresponding to different training tasks or the inference device;

[0299] The number of external iterations corresponding to different training tasks or the inference device.

[0300] In this embodiment, the first device can be a network-side device and the second device can be a terminal; or, the first device can be a network-side device and the second device can be a network-side device, such as in a scenario where multiple network-side devices aggregate training data into one network-side device for training; or, the first device can be a terminal and the second device can be a terminal, such as in a scenario where multiple terminals aggregate training data into one terminal for training.

[0301] Additionally, the candidate second device can be a network-side device or a terminal; the inference device can be a network-side device or a terminal. This application also provides a data collection apparatus, including:

[0302] The receiving module is configured to receive a first instruction from the first device, wherein the first instruction is used to instruct the second device to collect and report training data for training a specific AI model.

[0303] The processing module is used to collect training data and report the training data to the first device.

[0304] In some embodiments, the receiving module is used to receive the first instruction unicast by the first device, and the second device is a second device selected by the first device from candidate second devices according to a preset first filtering condition; or

[0305] The first device receives a first instruction broadcast to a candidate second device, the first instruction carrying a second filtering condition, the second filtering condition being used to filter the second device that reports the training data, and the second device meeting the second filtering condition.

[0306] In some embodiments, the processing module is configured to, if the second device receives the first instruction unicast by the first device, the second device collects and reports the training data; or

[0307] If the second device receives the first instruction broadcast by the first device, the second device collects and reports the training data.

[0308] In some embodiments, the candidate second device reports first training data and / or first parameters to the first device, whereby the first parameters may be judgment parameters of the first screening condition.

[0309] In some embodiments, the candidate second device reports only the first training data to the first device, and the first training data is used to determine the first parameter.

[0310] In some embodiments, the first parameter includes at least one of the following:

[0311] The data type of the candidate second device;

[0312] The data distribution parameters of the candidate second device;

[0313] The service type of the candidate second device;

[0314] The working scenario of the candidate second device;

[0315] The communication network access method of the candidate second device;

[0316] The channel quality of the candidate second device;

[0317] The degree of difficulty in collecting data by the candidate second device;

[0318] The power status of the candidate second device;

[0319] The storage state of the candidate second device.

[0320] In some embodiments, the processing module is further configured to send a first request to the first device, requesting the collection and reporting of training data.

[0321] In some embodiments, the first instruction of unicast includes at least one of the following:

[0322] The number of training data samples collected by the second device may be different or the same for different second devices;

[0323] The time it takes for the second device to collect training data may vary or be the same for different second devices.

[0324] The time at which the second device reports training data to the first device may vary or be the same for different second devices.

[0325] Does the collected data need to be preprocessed?

[0326] Methods for preprocessing the collected data;

[0327] The data format of the training data reported by the second device to the first device.

[0328] In some embodiments, the first instruction broadcast includes at least one of the following:

[0329] Identification of the candidate second device for data collection;

[0330] Identifier of a candidate second device that does not collect data;

[0331] The number of training data samples that the candidate second device needs to collect for data collection; the number of training data samples collected by different candidate second devices may be different or the same.

[0332] The time taken for candidate second devices to collect training data; the time taken for different candidate second devices to collect training data may be different or the same.

[0333] The time at which the candidate second device for data collection reports training data to the first device may vary or be the same for different candidate second devices.

[0334] Does the collected data need to be preprocessed?

[0335] Methods for preprocessing the collected data;

[0336] The data format of the training data reported by the candidate second device for data collection to the first device;

[0337] The first filtering condition.

[0338] In some embodiments, the inference device receives the trained AI model and hyperparameters sent by the first device.

[0339] In some embodiments, the AI ​​model is a meta-learning model, and the hyperparameters are determined by the first parameter.

[0340] In some embodiments, the hyperparameters include at least one of the following:

[0341] External learning rate;

[0342] The internal iterative learning rate corresponding to different training tasks or the inference device;

[0343] Meta-learning rate;

[0344] The number of internal iterations corresponding to different training tasks or the inference device;

[0345] The number of external iterations corresponding to different training tasks or the inference device.

[0346] In some embodiments, the data collection device further includes:

[0347] The inference module is used to perform performance verification on the AI ​​model; if the performance verification result meets the preset first condition, the AI ​​model is used for inference.

[0348] In some embodiments, the AI ​​model used for performance verification is the AI ​​model issued by the first device, or the AI ​​model issued by the first device after fine-tuning.

[0349] In this embodiment, the first device can be a network-side device and the second device can be a terminal; or, the first device can be a network-side device and the second device can be a network-side device, such as in a scenario where multiple network-side devices aggregate training data into one network-side device for training; or, the first device can be a terminal and the second device can be a terminal, such as in a scenario where multiple terminals aggregate training data into one terminal for training.

[0350] Additionally, the candidate second device can be a network-side device or a terminal; the inference device can also be a network-side device or a terminal. The data collection device in this application embodiment can be an electronic device, such as an electronic device with an operating system, or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the terminal can include, but is not limited to, the type of terminal 11 listed above; other devices can be servers, network attached storage (NAS), etc., and this application embodiment does not specifically limit the types.

[0351] The data collection device provided in this application embodiment can achieve... Figures 4 to 5The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0352] Optional, such as Figure 6 As shown, this application embodiment also provides a communication device 600, including a processor 601 and a memory 602. The memory 602 stores a program or instructions that can run on the processor 601. For example, when the communication device 600 is a first device, when the program or instructions are executed by the processor 601, they implement the various steps of the above-described data collection method embodiment and achieve the same technical effect. When the communication device 600 is a second device, when the program or instructions are executed by the processor 601, they implement the various steps of the above-described data collection method embodiment and achieve the same technical effect. To avoid repetition, further details are omitted here.

[0353] This application also provides a first device, which includes a processor and a memory. The memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, they implement the steps of the data collection method described above.

[0354] This application embodiment also provides a first device, including a processor and a communication interface, wherein the communication interface is used to send a first instruction to a second device, instructing the second device to collect and report training data for training a specific AI model; and to receive the training data reported by the second device; the processor is used to construct a dataset using the training data and train the specific AI model.

[0355] This application also provides a second device, which includes a processor and a memory. The memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, they implement the steps of the data collection method described above.

[0356] This application embodiment also provides a second device, including a processor and a communication interface, wherein the communication interface is used to receive a first instruction from a first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model; the processor is used to collect the training data and report the training data to the first device.

[0357] The first device mentioned above can be a network-side device or a terminal, and the second device can be a network-side device or a terminal.

[0358] When the first device and / or the second device are terminals, this application embodiment also provides a terminal, including a processor and a communication interface. This terminal embodiment corresponds to the above-described terminal-side method embodiment. All implementation processes and methods of the above-described method embodiments can be applied to this terminal embodiment and achieve the same technical effect. Specifically, Figure 7 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.

[0359] The terminal 700 includes, but is not limited to, at least some of the following components: radio frequency unit 701, network module 702, audio output unit 703, input unit 704, sensor 705, display unit 706, user input unit 707, interface unit 708, memory 709, and processor 710.

[0360] Those skilled in the art will understand that the terminal 700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0361] It should be understood that, in this embodiment, the input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0362] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 701 can transmit it to the processor 710 for processing; in addition, the radio frequency unit 701 can send uplink data to the network-side device. Typically, the radio frequency unit 701 includes, but is not limited to, an antenna, amplifier, transceiver, coupler, low-noise amplifier, duplexer, etc.

[0363] The memory 709 can be used to store software programs or instructions, as well as various data. The memory 709 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 709 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 709 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0364] Processor 710 may include one or more processing units; optionally, processor 710 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 710.

[0365] In some embodiments, the first device is a terminal, and the processor 710 is used to send a first instruction to the second device, instructing the second device to collect and report training data for training a specific AI model; receive the training data reported by the second device; construct a dataset using the training data; and train the specific AI model.

[0366] In some embodiments, the processor 710 is specifically configured to select N second devices from M candidate second devices according to a preset first filtering condition, and unicast the first instruction to the N second devices, where M and N are positive integers, and N is less than or equal to M; or

[0367] The first instruction is broadcast to the M candidate second devices. The first instruction carries a second filtering condition. The second filtering condition is used to filter the second devices that report the training data. The second devices meet the second filtering condition.

[0368] In some embodiments, the processor 710 is further configured to receive first training data and / or first parameters reported by the candidate second device, wherein the first parameters may be judgment parameters of the first screening condition.

[0369] In some embodiments, the processor 710 is configured to receive only the first training data reported by the candidate second device and determine the first parameter based on the first training data.

[0370] In some embodiments, the first parameter includes at least one of the following:

[0371] The data type of the candidate second device;

[0372] The data distribution parameters of the candidate second device;

[0373] The service type of the candidate second device;

[0374] The working scenario of the candidate second device;

[0375] The communication network access method of the candidate second device;

[0376] The channel quality of the candidate second device;

[0377] The degree of difficulty in collecting data by the candidate second device;

[0378] The power status of the candidate second device;

[0379] The storage state of the candidate second device.

[0380] In some embodiments, the first instruction of unicast includes at least one of the following:

[0381] The number of training data samples collected by the second device may be different or the same for different second devices;

[0382] The time it takes for the second device to collect training data may vary or be the same for different second devices.

[0383] The time at which the second device reports training data to the first device may vary or be the same for different second devices.

[0384] Does the collected data need to be preprocessed?

[0385] Methods for preprocessing the collected data;

[0386] The data format of the training data reported by the second device to the first device.

[0387] In some embodiments, the first instruction broadcast includes at least one of the following:

[0388] Identification of the candidate second device for data collection;

[0389] Identifier of a candidate second device that does not collect data;

[0390] The number of training data samples that the candidate second device needs to collect for data collection; the number of training data samples collected by different candidate second devices may be different or the same.

[0391] The time taken for candidate second devices to collect training data; the time taken for different candidate second devices to collect training data may be different or the same.

[0392] The time at which the candidate second device for data collection reports training data to the first device may vary or be the same for different candidate second devices.

[0393] Does the collected data need to be preprocessed?

[0394] Methods for preprocessing the collected data;

[0395] The data format of the training data reported by the candidate second device for data collection to the first device;

[0396] The first filtering condition.

[0397] In some embodiments, the processor 710 is also configured to send trained AI models and hyperparameters to L inference devices, wherein L is greater than M, equal to M, or less than M.

[0398] In some embodiments, the AI ​​model is a meta-learning model, and the hyperparameters are determined by the first parameter.

[0399] In some embodiments, the hyperparameters include at least one of the following:

[0400] External learning rate;

[0401] The internal iterative learning rate corresponding to different training tasks or the inference device;

[0402] Meta-learning rate;

[0403] The number of internal iterations corresponding to different training tasks or the inference device;

[0404] The number of external iterations corresponding to different training tasks or the inference device.

[0405] In this embodiment, the first device can be a network-side device and the second device can be a terminal; or, the first device can be a network-side device and the second device can be a network-side device, such as in a scenario where multiple network-side devices aggregate training data into one network-side device for training; or, the first device can be a terminal and the second device can be a terminal, such as in a scenario where multiple terminals aggregate training data into one terminal for training.

[0406] Additionally, the candidate second device can be a network-side device or a terminal; the inference device can be a network-side device or a terminal. In some embodiments, the second device is a terminal, and the processor 710 is used to receive a first instruction from the first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model; collect the training data, and report the training data to the first device.

[0407] In some embodiments, the processor 710 is configured to receive the first instruction unicast by the first device, wherein the second device is a second device selected by the first device from candidate second devices according to a preset first filtering condition; or

[0408] The first device receives a first instruction broadcast to a candidate second device, the first instruction carrying a second filtering condition, the second filtering condition being used to filter the second device that reports the training data, and the second device meeting the second filtering condition.

[0409] In some embodiments, the processor 710 is configured to collect and report the training data if the second device receives the first instruction unicast by the first device; or

[0410] If the second device receives the first instruction broadcast by the first device, the second device collects and reports the training data.

[0411] In some embodiments, the candidate second device reports first training data and / or first parameters to the first device, whereby the first parameters may be judgment parameters of the first screening condition.

[0412] In some embodiments, the candidate second device reports only the first training data to the first device, and the first training data is used to determine the first parameter.

[0413] In some embodiments, the first parameter includes at least one of the following:

[0414] The data type of the candidate second device;

[0415] The data distribution parameters of the candidate second device;

[0416] The service type of the candidate second device;

[0417] The working scenario of the candidate second device;

[0418] The communication network access method of the candidate second device;

[0419] The channel quality of the candidate second device;

[0420] The degree of difficulty in collecting data by the candidate second device;

[0421] The power status of the candidate second device;

[0422] The storage state of the candidate second device.

[0423] In some embodiments, the processor 710 is also configured to send a first request to the first device, requesting the collection and reporting of training data.

[0424] In some embodiments, the first instruction of unicast includes at least one of the following:

[0425] The number of training data samples collected by the second device may be different or the same for different second devices;

[0426] The time it takes for the second device to collect training data may vary or be the same for different second devices.

[0427] The time at which the second device reports training data to the first device may vary or be the same for different second devices.

[0428] Does the collected data need to be preprocessed?

[0429] Methods for preprocessing the collected data;

[0430] The data format of the training data reported by the second device to the first device.

[0431] In some embodiments, the first instruction broadcast includes at least one of the following:

[0432] Identification of the candidate second device for data collection;

[0433] Identifier of a candidate second device that does not collect data;

[0434] The number of training data samples that the candidate second device needs to collect for data collection; the number of training data samples collected by different candidate second devices may be different or the same.

[0435] The time taken for candidate second devices to collect training data; the time taken for different candidate second devices to collect training data may be different or the same.

[0436] The time at which the candidate second device for data collection reports training data to the first device may vary or be the same for different candidate second devices.

[0437] Does the collected data need to be preprocessed?

[0438] Methods for preprocessing the collected data;

[0439] The data format of the training data reported by the candidate second device for data collection to the first device;

[0440] The first filtering condition.

[0441] In some embodiments, the inference device receives the trained AI model and hyperparameters sent by the first device.

[0442] In some embodiments, the AI ​​model is a meta-learning model, and the hyperparameters are determined by the first parameter.

[0443] In some embodiments, the hyperparameters include at least one of the following:

[0444] External learning rate;

[0445] The internal iterative learning rate corresponding to different training tasks or the inference device;

[0446] Meta-learning rate;

[0447] The number of internal iterations corresponding to different training tasks or the inference device;

[0448] The number of external iterations corresponding to different training tasks or the inference device.

[0449] In some embodiments, the processor 710 is used to perform performance verification on the AI ​​model; if the performance verification result meets a preset first condition, the AI ​​model is used for inference.

[0450] In some embodiments, the AI ​​model used for performance verification is the AI ​​model issued by the first device, or the AI ​​model issued by the first device after fine-tuning.

[0451] When the first device and / or the second device are network-side devices, this application embodiment also provides a network-side device, including a processor and a communication interface. This network-side device embodiment corresponds to the above-described network-side device method embodiment. All implementation processes and methods of the above-described method embodiments can be applied to this network-side device embodiment and achieve the same technical effects.

[0452] Specifically, embodiments of this application also provide a network-side device. For example... Figure 8 As shown, the network-side device 800 includes: an antenna 81, a radio frequency (RF) device 82, a baseband device 83, a processor 84, and a memory 85. The antenna 81 is connected to the RF device 82. In the uplink direction, the RF device 82 receives information through the antenna 81 and transmits the received information to the baseband device 83 for processing. In the downlink direction, the baseband device 83 processes the information to be transmitted and sends it to the RF device 82. The RF device 82 processes the received information and transmits it through the antenna 81.

[0453] The method executed by the network-side device in the above embodiments can be implemented in the baseband device 83, which includes a baseband processor.

[0454] Baseband device 83 may include, for example, at least one baseband board on which multiple chips are disposed, such as... Figure 8 As shown, one of the chips is, for example, a baseband processor, which is connected to the memory 85 via a bus interface to call the program in the memory 85 and execute the network device operation shown in the above method embodiment.

[0455] The network-side device may also include a network interface 86, such as a common public radio interface (CPRI).

[0456] Specifically, the network-side device 800 of this embodiment of the invention further includes: instructions or programs stored in memory 85 and executable on processor 84. Processor 84 calls the instructions or programs in memory 85 to execute the data collection method described above and achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0457] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described data collection method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0458] The processor is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0459] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above data collection method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0460] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0461] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described data collection method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0462] This application also provides a data collection system, including a first device and a second device, wherein the first device can be used to perform the steps of the data collection method described above, and the second device can be used to perform the steps of the data collection method described above.

[0463] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0464] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0465] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A data collection method, characterized in that, include: The first device sends a first instruction to the second device, instructing the second device to collect and report training data used for training a specific AI model; The first device receives the training data reported by the second device; The first device uses the training data to construct a dataset and train the specific AI model. The first device sending a first instruction to the second device includes: The first device selects N second devices from M candidate second devices according to a preset first filtering condition, and unicasts the first instruction to the N second devices, where M and N are positive integers, and N is less than or equal to M; or The first device broadcasts the first instruction to the M candidate second devices. The first instruction carries a second filtering condition, which is used to filter the second devices that report the training data. The second devices meet the second filtering condition.

2. The data collection method according to claim 1, characterized in that, Before the first device sends the first instruction to the second device, the method further includes: The first device receives the first training data and / or the first parameter reported by the candidate second device, wherein the first parameter is the judgment parameter of the first screening condition.

3. The data collection method according to claim 2, characterized in that, The first device receives only the first training data reported by the candidate second device and determines the first parameter based on the first training data.

4. The data collection method according to claim 2 or 3, characterized in that, The first parameter includes at least one of the following: The data type of the candidate second device; The data distribution parameters of the candidate second device; The service type of the candidate second device; The working scenario of the candidate second device; The communication network access method of the candidate second device; The channel quality of the candidate second device; The degree of difficulty in collecting data by the candidate second device; The power status of the candidate second device; The storage state of the candidate second device.

5. The data collection method according to claim 1, characterized in that, The first instruction for unicast includes at least one of the following: The number of training data samples collected by the second device; The time it takes for the second device to collect training data; The time it takes for the second device to report training data to the first device; Does the collected data need to be preprocessed? Methods for preprocessing the collected data; The data format of the training data reported by the second device to the first device.

6. The data collection method according to claim 1, characterized in that, The first instruction broadcast includes at least one of the following: Identification of the candidate second device for data collection; Identifier of a candidate second device that does not collect data; The number of training data samples that the candidate second device needs to collect for data collection; The time required for the candidate second device to collect training data; The time when the candidate second device for data collection reports training data to the first device; Does the collected data need to be preprocessed? Methods for preprocessing the collected data; The data format of the training data reported by the candidate second device for data collection to the first device; The first filtering condition.

7. The data collection method according to claim 2, characterized in that, After training a specific AI model, the method further includes: The first device sends the trained AI model and hyperparameters to L inference devices, where L is greater than M, equal to M, or less than M.

8. The data collection method according to claim 7, characterized in that, The AI ​​model is a meta-learning model, and the hyperparameters are determined by the first parameter.

9. The data collection method according to claim 7, characterized in that, The hyperparameters include at least one of the following: External learning rate; The internal iterative learning rate corresponding to different training tasks or the inference device; Meta-learning rate; The number of internal iterations corresponding to different training tasks or the inference device; The number of external iterations corresponding to different training tasks or the inference device.

10. The data collection method according to claim 1, characterized in that, The first device is a network-side device, and the second device is a terminal; or The first device is a network-side device, and the second device is a network-side device; or The first device is a terminal, and the second device is a terminal.

11. A data collection method, characterized in that, include: The second device receives a first instruction from the first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model; The second device collects training data and reports the training data to the first device; Wherein, the second device receiving the first instruction from the first device includes: The second device receives the first instruction unicast by the first device, and the second device is the second device selected by the first device from the candidate second devices according to a preset first filtering condition; or The second device receives the first instruction broadcast by the first device to the candidate second device. The first instruction carries a second filtering condition, which is used to filter the second devices that report the training data. The second device meets the second filtering condition.

12. The data collection method according to claim 11, characterized in that, The second device collects training data and reports the training data to the first device, including: If the second device receives the first instruction unicast by the first device, the second device collects and reports the training data; or If the second device receives the first instruction broadcast by the first device, the second device collects and reports the training data.

13. The data collection method according to claim 11, characterized in that, Before the second device receives the first instruction from the first device, the method further includes: The candidate second device reports the first training data and / or the first parameter to the first device, where the first parameter is the judgment parameter of the first screening condition.

14. The data collection method according to claim 13, characterized in that, The candidate second device reports only the first training data to the first device, and the first training data is used to determine the first parameter.

15. The data collection method according to claim 13, characterized in that, The first parameter includes at least one of the following: The data type of the candidate second device; The data distribution parameters of the candidate second device; The service type of the candidate second device; The working scenario of the candidate second device; The communication network access method of the candidate second device; The channel quality of the candidate second device; The degree of difficulty in collecting data by the candidate second device; The power status of the candidate second device; The storage state of the candidate second device.

16. The data collection method according to claim 11, characterized in that, Before reporting the training data to the first device, the method further includes: The second device sends a first request to the first device, requesting the collection and reporting of training data.

17. The data collection method according to claim 11, characterized in that, The first instruction for unicast includes at least one of the following: The number of training data samples collected by the second device; The time it takes for the second device to collect training data; The time it takes for the second device to report training data to the first device; Does the collected data need to be preprocessed? Methods for preprocessing the collected data; The data format of the training data reported by the second device to the first device.

18. The data collection method according to claim 11, characterized in that, The first instruction broadcast includes at least one of the following: Identification of the candidate second device for data collection; Identifier of a candidate second device that does not collect data; The number of training data samples that the candidate second device needs to collect for data collection; The time required for the candidate second device to collect training data; The time when the candidate second device for data collection reports training data to the first device; Does the collected data need to be preprocessed? Methods for preprocessing the collected data; The data format of the training data reported by the candidate second device for data collection to the first device; The first filtering condition.

19. The data collection method according to claim 13, characterized in that, After the second device collects training data and reports the training data to the first device, the method further includes: The inference device receives the trained AI model and hyperparameters sent by the first device.

20. The data collection method according to claim 19, characterized in that, The AI ​​model is a meta-learning model, and the hyperparameters are determined by the first parameter.

21. The data collection method according to claim 19, characterized in that, The hyperparameters include at least one of the following: External learning rate; The internal iterative learning rate corresponding to different training tasks or the inference device; Meta-learning rate; The number of internal iterations corresponding to different training tasks or the inference device; The number of external iterations corresponding to different training tasks or the inference device.

22. The data collection method according to claim 19, characterized in that, After the inference device receives the trained AI model and hyperparameters sent by the first device, the method further includes: The inference device performs performance verification on the AI ​​model; If the performance verification result meets the preset first condition, the inference device will use the AI ​​model for inference.

23. The data collection method according to claim 22, characterized in that, The AI ​​model used for performance verification is either the AI ​​model issued by the first device, or a fine-tuned version of the AI ​​model issued by the first device.

24. The data collection method according to claim 11, characterized in that, The first device is a network-side device, and the second device is a terminal; or The first device is a network-side device, and the second device is a network-side device; or The first device is a terminal, and the second device is a terminal.

25. A data collection device, characterized in that, include: The sending module is used to send a first instruction to the second device, instructing the second device to collect and report training data for training a specific AI model; A receiving module is used to receive training data reported by the second device; The training module is used to construct a dataset using the training data and train the specific AI model. Specifically, the sending module is used to select N second devices from M candidate second devices according to a preset first filtering condition, and unicast the first instruction to the N second devices, where M and N are positive integers, and N is less than or equal to M; or The first instruction is broadcast to the M candidate second devices. The first instruction carries a second filtering condition. The second filtering condition is used to filter the second devices that report the training data. The second devices meet the second filtering condition.

26. A data collection device, characterized in that, include: The receiving module is used to receive a first instruction from the first device, the first instruction being used to instruct the second device to collect and report training data for training a specific AI model. The processing module is used to collect training data and report the training data to the first device; Specifically, the receiving module is used to receive the first instruction unicast by the first device, and the second device is the second device selected by the first device from the candidate second devices according to the preset first filtering conditions; or The first device receives a first instruction broadcast to a candidate second device, the first instruction carrying a second filtering condition, the second filtering condition being used to filter the second device that reports the training data, and the second device meeting the second filtering condition.

27. The data collection apparatus according to claim 26, characterized in that, The data collection device also includes: The inference module is used to perform performance verification on the AI ​​model; if the performance verification result meets the preset first condition, the AI ​​model is used for inference.

28. A first device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the data collection method as described in any one of claims 1 to 10.

29. A second device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the data collection method as described in any one of claims 11 to 24.

30. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the data collection method as described in any one of claims 1-10, or implement the steps of the data collection method as described in any one of claims 11-24.