Model data acquisition method, apparatus and system

By generating synthetic data that meets diversity requirements on the terminal device, the problem of limited data on the device is solved, the quality of model training and update is improved, and the effective deployment and inference performance of the model on the terminal device is ensured.

WO2025140663A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/143477
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-28
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In terminal devices or specific scenarios, the measured data obtained by the device side is limited, resulting in low quality of model training data, making it difficult to ensure the convergence performance and generalization of the model.

Method used

By receiving the original data and feature information from the second node, a synthetic data that meets the diversity requirements and has the same data distribution as the measured data, and data processing is performed using the first node or circuit to improve the quality of the synthetic data.

Benefits of technology

Improves the performance of model training or updates, ensuring the effective deployment and inference performance of the model on end devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143477_03072025_PF_FP_ABST
    Figure CN2024143477_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an artificial intelligence (AI) model data acquisition method, apparatus and a system. The method may comprise: receiving first information from a second node, the first information comprising the following types of information: original data and feature information of the original data; and on the basis of the first information, generating multiple pieces of synthetic data, the multiple pieces of synthetic data being used for model processing. In said example, synthetic data that meets diversity requirements and has the same data distribution as measured data can be generated on the basis of original data and feature information of the original data, such that the quality of the synthetic data is improved, thereby improving the performance of model training or updating.
Need to check novelty before this filing date? Find Prior Art

Description

Model data acquisition method, device and system

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 29, 2023, with application number 202311871100.7, and priority to the Chinese patent application entitled “Model Data Acquisition Method, Device and System”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI), and in particular to a method, device, and system for acquiring model data. Background Art

[0003] With the advancement of artificial intelligence research, the application scenarios of neural networks continue to expand, including smart healthcare and intelligent networks. Therefore, how to effectively train neural networks for specific scenarios has become an important research direction. The three key factors in neural network training are data, neural networks, and computing platforms. With the widespread deployment of graphics processing units (GPUs), the proportion of devices equipped with neural network computing platforms has gradually increased, resulting in an increasing demand for specialized neural network training for workers.

[0004] The quality of training data is a key factor influencing model training and fine-tuning performance. However, acquiring training data for training or fine-tuning specialized models or scenario-specific models deployed on discrete devices, such as those deployed on the terminal side or at specific base stations, is highly challenging. This is partly because model training or fine-tuning requires sufficient training data from multiple scenarios to ensure model convergence, but the measured data available on the device side is limited, making it difficult to ensure model performance using only data collected by the device. Furthermore, model training or fine-tuning requires sufficiently rich data features to ensure model generalization and robustness, but the measured data available on the device side is limited, resulting in limited data feature diversity.

[0005] In existing technologies, a central node receives and merges raw local data from multiple devices as a training set to increase the amount and diversity of training data. The central node then uses this merged training set to train or update the model. The trained model is deployed on the device for model inference. However, the central node generates synthetic data based solely on the raw local data from multiple devices as input to the generative model, resulting in low-quality synthetic data. Summary of the Invention

[0006] The present application discloses a model data acquisition method, device and system, which can improve the quality of synthetic data.

[0007] In a first aspect, embodiments of the present application provide a model data acquisition method, performed by a first node or a circuit for the first node. The method includes: receiving first information from a second node. The first information includes the following types of information: original data and feature information of the original data; and generating multiple synthetic data based on the first information. The multiple synthetic data are used for model processing.

[0008] In an embodiment of the present application, the first node or a circuit used for the first node, such as a chip, generates a plurality of synthetic data based on the acquired original data and characteristic information of the original data.

[0009] Based on the original data and its feature information, synthetic data that meets diversity requirements and has the same data distribution as the measured data can be generated, improving the quality of the synthetic data and, in turn, helping to improve the performance of model training or updates. The following description uses the first node as the execution entity. It is understood that the first node can also be replaced by a circuit for the first node, such as a chip.

[0010] In one possible implementation, the first node further sends second information to the second node. The second information indicates a processing method for the second node to obtain the first information. The processing method for the second node to obtain the first information includes a method for acquiring feature information of the original data.

[0011] Based on the instruction of the second information, the second node processes the original data to obtain characteristic information of the original data.

[0012] In a possible implementation, the first information further includes at least one of the following types of information: a priority of the original data, and a data tag of the original data.

[0013] By reporting the priority of the original data, the first node determines the number of similar synthetic data to be generated according to the priority or probability of each original data, so that the distribution of the synthetic data is consistent with expectations.

[0014] In a possible implementation, the processing method for the second node to obtain the first information also includes a method for obtaining the data label of the original data.

[0015] Based on the method for obtaining the data label of the original data, the second node processes the original data to obtain the data label of the original data.

[0016] In a possible implementation manner, the first node further sends third information to the second node, where the third information indicates a reporting configuration of the first information.

[0017] The first node sends the third information to the second node, so that the second node reports the first information based on the reporting configuration of the first information indicated in the third information. This allows the separate data provider and the data synthesizer to align their understanding of the data, thereby improving the quality of the synthesized data.

[0018] Optionally, the reporting configuration includes at least one of the following: the type of the first information, and the data volume of the first information.

[0019] The first node sends a reporting configuration of the type of first information and / or a reporting configuration of the data amount of the first information to the second node, so that the second node sends the corresponding type of first information or the corresponding data amount of first information to the first node, which helps to improve the quality of the synthesized data.

[0020] In one possible implementation, the first node further sends fourth information to the second node. The fourth information indicates that the first node has the ability to jointly process feature information of multiple raw data in the same group. This joint processing may include, for example, performing a weighted summation of the feature information of multiple raw data in the same group.

[0021] The first node sends the fourth information to the second node so that the second node can divide the original data that can be jointly processed into the same group, while dividing the original data that cannot be jointly processed into the same group, thereby avoiding the fusion of original data from different groups to generate synthetic data that does not meet expectations.

[0022] In a possible implementation manner, the first information further includes the following type of information: group information of the original data.

[0023] The second node reports the group information of the original data to the first node, and then the first node can jointly process the feature information of multiple original data in the same group.

[0024] In one possible implementation, the first node further sends at least one subset of the plurality of synthesized data to the second node. The first node further receives fifth information from the second node. The fifth information indicates an evaluation result of at least one synthesized data in the at least one subset. The evaluation result indicates at least one synthesized data in the at least one subset that needs to be removed, or indicates at least one synthesized data in the at least one subset that does not need to be removed.

[0025] Optionally, the evaluation result is used to filter the plurality of synthetic data to obtain at least one filtered synthetic data, wherein the at least one filtered synthetic data is used for processing the model.

[0026] In one possible implementation, the at least one synthesized data after screening is determined based on the distance between the first synthesized data to be eliminated and other synthesized data in the multiple synthesized data, and the first synthesized data to be eliminated is the synthesized data that needs to be eliminated determined based on the fifth information.

[0027] Screening synthetic data through the synthetic data evaluation process can effectively improve the quality of synthetic data.

[0028] Alternatively, an embodiment of the present application provides a model data acquisition method, which is performed by a first node or a circuit for the first node. The method includes: the first node receiving sixth information from a second node. The sixth information includes the following types of information: original data. The first node generates multiple synthetic data based on the sixth information and feature information of the original data. The multiple synthetic data are used for model processing. The following description uses the first node as the execution entity, but it is understandable that the first node can also be replaced by a circuit for the first node, such as a chip.

[0029] In a first possible implementation manner, the first node further receives eleventh information from the third node, where the eleventh information includes the following type of information: characteristic information of the original data.

[0030] In a second possible implementation, the first node further receives ninth information from the second node, the ninth information indicating a method by which the first node obtains the characteristic information of the original data. Furthermore, the first node processes the original data based on the ninth information to obtain the characteristic information of the original data.

[0031] In a third possible implementation, the first node further receives ninth information from the third node, the ninth information indicating a method by which the first node obtains the characteristic information of the original data. Furthermore, the first node processes the original data based on the ninth information to obtain the characteristic information of the original data.

[0032] In a fourth possible implementation manner, the characteristic information of the original data is determined based on a method for obtaining the characteristic information of the original data, and the method for obtaining the characteristic information of the original data is determined by the first node.

[0033] In a possible implementation, the sixth information further includes at least one of the following types of information: the priority of the original data, or a data tag of the original data.

[0034] By reporting the priority of the original data, the first node determines the number of similar synthetic data to be generated according to the priority or probability of each original data, so that the distribution of the synthetic data is consistent with expectations.

[0035] In one possible implementation, the first node also sends eighth information to the second node, where the eighth information indicates how the second node processes the sixth information, and the processing method for the second node to process the sixth information includes a method for obtaining a data label of the original data.

[0036] Based on the method for obtaining the data label of the original data, the second node processes the original data to obtain the data label of the original data.

[0037] In a possible implementation manner, the first node further sends seventh information to the second node, where the seventh information indicates a reporting configuration of the sixth information.

[0038] The first node sends the seventh information to the second node, so that the second node reports the sixth information based on the reporting configuration indicated in the seventh information. This allows the separate data provider and the data synthesizer to align their understanding of the data, thereby improving the quality of the synthesized data.

[0039] Optionally, the reporting configuration includes at least one of the following: the type of the sixth information, or the data volume of the sixth information.

[0040] The first node sends a reporting configuration of the type of the sixth information and / or a reporting configuration of the data amount of the sixth information to the second node, so that the second node sends the corresponding type of the sixth information, or the corresponding data amount of the sixth information to the first node, which helps to improve the quality of the synthesized data.

[0041] In one possible implementation, the first node further sends fourth information to the second node. The fourth information indicates that the first node has the ability to jointly process feature information of multiple raw data in the same group. This joint processing may include, for example, performing a weighted summation of the feature information of multiple raw data in the same group.

[0042] The first node sends the fourth information to the second node so that the second node can divide the original data that can be jointly processed into the same group, while dividing the original data that cannot be jointly processed into the same group, thereby avoiding the fusion of original data from different groups to generate synthetic data that does not meet expectations.

[0043] In a possible implementation manner, the sixth information further includes the following type of information: group information of the original data.

[0044] The second node reports the group information of the original data to the first node, and then the first node can jointly process the feature information of multiple original data in the same group.

[0045] In one possible implementation, the first node further sends at least one subset of the plurality of synthesized data to the second node. The first node further receives fifth information from the second node. The fifth information indicates an evaluation result of at least one synthesized data in the at least one subset. The evaluation result indicates at least one synthesized data in the at least one subset that needs to be removed, or indicates at least one synthesized data in the at least one subset that does not need to be removed.

[0046] Optionally, the evaluation result is used to filter the plurality of synthetic data to obtain at least one filtered synthetic data, wherein the at least one filtered synthetic data is used for processing the model.

[0047] In one possible implementation, the at least one synthesized data after screening is determined based on the distance between the first synthesized data to be eliminated and other synthesized data in the multiple synthesized data, and the first synthesized data to be eliminated is the synthesized data that needs to be eliminated determined based on the fifth information.

[0048] Screening synthetic data through the synthetic data evaluation process can effectively improve the quality of synthetic data.

[0049] In a second aspect, embodiments of the present application provide a model data acquisition method, performed by a second node or a circuit for the second node. The method includes: sending first information to a first node. The first information includes the following types of information: raw data and characteristic information of the raw data. The following description uses the second node as the execution subject, but it is understood that the second node can also be replaced by a circuit for the second node, such as a chip.

[0050] In an embodiment of the present application, the second node sends original data and characteristic information of the original data to the first node, and the original data and the characteristic information of the original data are used to generate multiple synthetic data.

[0051] Based on the original data and its feature information, synthetic data that meets diversity requirements and has the same data distribution as the measured data can be generated, which can improve the quality of the synthetic data and thus help improve the performance of model training or updating.

[0052] In one possible implementation, the second node further receives second information from the first node. The second information indicates a processing method used by the second node to obtain the first information. Exemplarily, the processing method used by the second node to obtain the first information includes a method for acquiring characteristic information of the original data. Furthermore, the second node obtains the first information based on the second information.

[0053] In a possible implementation, the first information further includes at least one of the following types of information: a data label of the original data, or a priority of the original data.

[0054] In a possible implementation, the processing method further includes a method for obtaining data labels of the original data.

[0055] In a possible implementation manner, the second node further receives third information from the first node, where the third information indicates a reporting configuration of the first information.

[0056] Optionally, the reporting configuration includes at least one of the following: the type of the first information, and the data volume of the first information.

[0057] In a possible implementation, the second node further receives fourth information from the first node, where the fourth information indicates that the first node has the ability to jointly process feature information of multiple original data in the same group.

[0058] In a possible implementation manner, the first information further includes the following type of information: group information of the original data.

[0059] In one possible implementation, the second node further receives at least one subset of the plurality of synthesized data from the first node. The second node further sends fifth information to the first node, the fifth information indicating an evaluation result of at least one synthesized data in the at least one subset. The evaluation result indicates at least one synthesized data in the at least one subset that needs to be eliminated, or indicates at least one synthesized data in the at least one subset that does not need to be eliminated.

[0060] In a possible implementation, the evaluation result is determined based on a distance between at least one synthetic data in the at least one subset and the local data.

[0061] In a possible implementation, the evaluation result is used to filter the multiple synthetic data to obtain at least one filtered synthetic data; wherein, the at least one filtered synthetic data is used for model processing.

[0062] Alternatively, an embodiment of the present application provides a model data acquisition method, performed by a second node or a circuit for the second node. The method includes: the second node sending sixth information to the first node. The sixth information includes the following types of information: original data.

[0063] The second node further sends ninth information to the first node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data.

[0064] In one possible implementation, the second node also receives eighth information from the first node, where the eighth information indicates how the second node processes the sixth information, and the processing method for the second node to process the sixth information includes a method for obtaining a data label of the original data.

[0065] In a possible implementation manner, the second node further receives seventh information from the first node, where the seventh information indicates a reporting configuration of the sixth information.

[0066] In a possible implementation, the second node further receives fourth information from the first node, where the fourth information indicates that the first node has the ability to jointly process feature information of multiple original data in the same group.

[0067] In one possible implementation, the second node further receives at least one subset of the plurality of synthesized data from the first node. The second node further sends fifth information to the first node, the fifth information indicating an evaluation result of at least one synthesized data in the at least one subset. The evaluation result indicates at least one synthesized data in the at least one subset that needs to be eliminated, or indicates at least one synthesized data in the at least one subset that does not need to be eliminated.

[0068] In a possible implementation, the evaluation result is determined based on a distance between at least one synthetic data in the at least one subset and the local data.

[0069] In a possible implementation, the evaluation result is used to filter the multiple synthetic data to obtain at least one filtered synthetic data; wherein, the at least one filtered synthetic data is used for model processing.

[0070] In a third aspect, embodiments of the present application provide a model data acquisition method, performed by a third node or a circuit for a third node. The method includes: sending eleventh information to a first node, the eleventh information including the following type of information: characteristic information of the original data. This allows the first node to obtain synthesized data based on the characteristic information of the original data. In other words, the characteristic information of the original data is used to generate the synthesized data.

[0071] Alternatively, an embodiment of the present application provides a model data acquisition method, performed by a third node or a circuit for a third node. The method includes: sending ninth information to a first node, the ninth information indicating a method for the first node to obtain characteristic information of the raw data. This allows the first node to process the raw data based on the ninth information to obtain the characteristic information of the raw data. In other words, the method for obtaining characteristic information of the raw data is used to obtain characteristic information of the raw data.

[0072] In a fourth aspect, the present application provides a model data acquisition device, which may include a transceiver module and a processing module, as follows:

[0073] a transceiver module, configured to receive first information from a second node, wherein the first information includes the following types of information: original data and characteristic information of the original data;

[0074] A processing module is used to generate a plurality of synthetic data based on the first information, and the plurality of synthetic data are used for processing the model.

[0075] In a possible implementation, the transceiver module is further used to send second information to the second node, where the second information indicates a processing method for the second node to obtain the first information, and the processing method for the second node to obtain the first information includes a method for obtaining characteristic information of the original data.

[0076] In a possible implementation manner, the first information further includes the following types of information: a data label of the original data and / or a priority of the original data.

[0077] In a possible implementation, the processing method for the second node to obtain the first information also includes a method for obtaining the data label of the original data.

[0078] In a possible implementation, the transceiver module is further configured to send third information to the second node, where the third information indicates a reporting configuration of the first information.

[0079] In a possible implementation manner, the reporting configuration includes at least one of the following: the type of the first information, and the data volume of the first information.

[0080] In a possible implementation, the transceiver module is further configured to send fourth information to the second node, where the fourth information indicates that the apparatus has the ability to jointly process feature information of multiple original data in the same group.

[0081] In a possible implementation manner, the first information further includes the following type of information: group information of the original data.

[0082] In a possible implementation, the transceiver module is further configured to send at least a subset of the plurality of synthesized data to the second node;

[0083] Receive fifth information from the second node, wherein the fifth information indicates an evaluation result of at least one synthetic data in the at least one subset, and the evaluation result indicates at least one synthetic data that needs to be eliminated in the at least one subset, or indicates at least one synthetic data that does not need to be eliminated in the at least one subset.

[0084] In a possible implementation, the evaluation result is used to filter the multiple synthetic data to obtain at least one filtered synthetic data; wherein the at least one filtered synthetic data is used for processing the model.

[0085] In one possible implementation, the at least one synthesized data after screening is determined based on the distance between the first synthesized data to be eliminated and other synthesized data in the multiple synthesized data, and the first synthesized data to be eliminated is the synthesized data that needs to be eliminated determined based on the fifth information.

[0086] Alternatively, the present application provides a model data acquisition device, which may include a transceiver module and a processing module, as follows:

[0087] The transceiver module is configured to receive sixth information from the second node, where the sixth information includes the following types of information: original data.

[0088] The processing module is configured to generate a plurality of synthetic data based on the sixth information and the characteristic information of the original data. The plurality of synthetic data are used for processing the model.

[0089] In a first possible implementation manner, the transceiver module is further configured to receive eleventh information from the third node, where the eleventh information includes the following type of information: characteristic information of the original data.

[0090] In a second possible implementation manner, the transceiver module is further configured to receive ninth information from the second node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data.

[0091] In a third possible implementation manner, the transceiver module is further configured to receive ninth information from the third node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data.

[0092] In a fourth possible implementation manner, the characteristic information of the original data is determined based on a method for obtaining the characteristic information of the original data, and the method for obtaining the characteristic information of the original data is determined by the first node.

[0093] In one possible implementation, the transceiver module is further used to send eighth information to the second node, where the eighth information indicates how the second node processes the sixth information, and the processing method for the second node to process the sixth information includes a method for obtaining a data tag of the original data.

[0094] In a possible implementation, the transceiver module is further configured to send seventh information to the second node, where the seventh information indicates a reporting configuration of the sixth information.

[0095] Optionally, the reporting configuration includes at least one of the following: the type of the sixth information, and the data volume of the sixth information.

[0096] In one possible implementation, the transceiver module is further configured to send fourth information to the second node. The fourth information indicates that the first node has the ability to jointly process feature information of multiple raw data in the same group. The joint processing may include, for example, performing a weighted summation of the feature information of the multiple raw data in the same group.

[0097] In a possible implementation manner, the sixth information further includes the following type of information: group information of the original data.

[0098] In one possible implementation, the transceiver module is further configured to send at least one subset of the plurality of synthesized data to the second node. The first node further receives fifth information from the second node. The fifth information indicates an evaluation result of at least one synthesized data in the at least one subset. The evaluation result indicates at least one synthesized data in the at least one subset that needs to be eliminated, or indicates at least one synthesized data in the at least one subset that does not need to be eliminated.

[0099] Optionally, the evaluation result is used to filter the plurality of synthetic data to obtain at least one filtered synthetic data, wherein the at least one filtered synthetic data is used for processing the model.

[0100] In one possible implementation, the at least one synthesized data after screening is determined based on the distance between the first synthesized data to be eliminated and other synthesized data in the multiple synthesized data, and the first synthesized data to be eliminated is the synthesized data that needs to be eliminated determined based on the fifth information.

[0101] In a fifth aspect, the present application provides a model data acquisition device, which may include a transceiver module, specifically as follows:

[0102] The transceiver module is configured to send first information to the first node, where the first information includes the following types of information: original data and characteristic information of the original data.

[0103] In a possible implementation, the transceiver module is further configured to receive second information from the first node, where the second information indicates a processing method for the apparatus to obtain the first information, the processing method including a method for acquiring characteristic information of the original data;

[0104] The method further includes a processing module configured to obtain the first information based on the second information.

[0105] In a possible implementation manner, the first information further includes the following types of information: a data label of the original data and / or a priority of the original data.

[0106] In a possible implementation, the processing method further includes a method for obtaining data labels of the original data.

[0107] In a possible implementation, the transceiver module is further configured to receive third information from the first node, where the third information indicates a reporting configuration of the first information.

[0108] In a possible implementation manner, the reporting configuration includes at least one of the following: the type of the first information, and the data volume of the first information.

[0109] In a possible implementation, the transceiver module is further configured to receive fourth information from the first node, where the fourth information indicates that the first node has the ability to jointly process feature information of multiple original data in the same group.

[0110] In a possible implementation manner, the first information further includes the following type of information: group information of the original data.

[0111] In a possible implementation, the transceiver module is further configured to receive at least a subset of the plurality of synthesized data from the first node;

[0112] Send fifth information to the first node, where the fifth information indicates an evaluation result of at least one synthetic data in the at least one subset, where the evaluation result indicates at least one synthetic data that needs to be eliminated in the at least one subset, or indicates at least one synthetic data that does not need to be eliminated in the at least one subset.

[0113] In a possible implementation manner, the evaluation result is determined based on a distance between at least one synthetic data in the at least one subset and the local data.

[0114] In a possible implementation, the evaluation result is used to filter the multiple synthetic data to obtain at least one filtered synthetic data; wherein, the at least one filtered synthetic data is used for model processing.

[0115] Alternatively, the present application provides a model data acquisition device, which may include a transceiver module, specifically as follows:

[0116] The transceiver module is configured to send sixth information to the first node. The sixth information includes the following types of information: original data.

[0117] The transceiver module is further configured to send ninth information to the first node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data.

[0118] In one possible implementation, the transceiver module is used to receive eighth information from the first node, where the eighth information indicates a processing method for the second node to obtain the sixth information, and the processing method for the second node to obtain the sixth information includes a method for obtaining a data tag of the original data.

[0119] In a possible implementation, the transceiver module is configured to receive seventh information from the first node, where the seventh information indicates a reporting configuration of the sixth information.

[0120] In a possible implementation, the transceiver module is configured to receive fourth information from the first node, where the fourth information indicates that the first node has the ability to jointly process feature information of multiple original data in the same group.

[0121] In one possible implementation, the transceiver module is configured to receive at least one subset of the plurality of synthesized data from the first node. The second node further transmits fifth information to the first node, the fifth information indicating an evaluation result of at least one synthesized data in the at least one subset. The evaluation result indicates at least one synthesized data in the at least one subset that needs to be eliminated, or indicates at least one synthesized data in the at least one subset that does not need to be eliminated.

[0122] In a possible implementation, the evaluation result is determined based on a distance between at least one synthetic data in the at least one subset and the local data.

[0123] In a possible implementation, the evaluation result is used to filter the multiple synthetic data to obtain at least one filtered synthetic data; wherein, the at least one filtered synthetic data is used for model processing.

[0124] In a sixth aspect, the present application provides a model data acquisition device, which may include a transceiver module, wherein:

[0125] The transceiver module is configured to send eleventh information to the first node, where the eleventh information includes the following types of information: characteristic information of original data.

[0126] Alternatively, the present application provides a model data acquisition device, which may include a transceiver module, wherein:

[0127] The transceiver module is configured to send ninth information to the first node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data.

[0128] In a seventh aspect, the present application provides a model data acquisition device, comprising a processor and a memory; wherein the memory is used to store program code, and the processor is used to call the program code to execute the method provided in any possible implementation manner of the first aspect.

[0129] In an eighth aspect, the present application provides a model data acquisition device, comprising a processing circuit and a memory; wherein the memory is used to store program code, and the processing circuit is used to call the program code to execute the method provided in any possible implementation manner of the second aspect.

[0130] In a ninth aspect, the present application provides a model data acquisition device, comprising a processing circuit and a memory; wherein the memory is used to store program code, and the processing circuit is used to call the program code to execute a method provided in any possible implementation manner of the third aspect.

[0131] In a tenth aspect, the present application provides a model data acquisition system, comprising an apparatus as provided in any possible implementation of the fourth aspect, and an apparatus as provided in any possible implementation of the fifth aspect; or comprising an apparatus as provided in any possible implementation of the seventh aspect, and an apparatus as provided in any possible implementation of the eighth aspect; or comprising an apparatus as provided in any possible implementation of the fourth aspect, and an apparatus as provided in any possible implementation of the fifth aspect, and an apparatus as provided in any possible implementation of the sixth aspect; or comprising an apparatus as provided in any possible implementation of the seventh aspect, and an apparatus as provided in any possible implementation of the eighth aspect, and an apparatus as provided in any possible implementation of the ninth aspect.

[0132] In the eleventh aspect, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method provided in any possible implementation of the first aspect, or the method provided in any possible implementation of the second aspect, or the method provided in any possible implementation of the third aspect.

[0133] In the twelfth aspect, the present application provides a computer program product, characterized in that when the computer program product is run on a computer, the computer is caused to execute the method provided in any possible implementation of the first aspect, or the method provided in any possible implementation of the second aspect, or the method provided in any possible implementation of the third aspect.

[0134] It is understandable that the apparatus described in the fourth aspect to the apparatus described in the ninth aspect, the system described in the tenth aspect, the computer-readable storage medium described in the eleventh aspect, or the computer program product described in the twelfth aspect are all used to execute any of the methods provided in the first aspect, any of the methods provided in the second aspect, or any of the methods provided in the third aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0135] The following is an introduction to the drawings used in the embodiments of this application.

[0136] FIG1 is a simplified schematic diagram of a wireless communication system provided by an embodiment of the present application;

[0137] FIG2a is a schematic diagram of a communication system provided by an embodiment of the present application;

[0138] FIG2b is a schematic diagram of another communication system provided in an embodiment of the present application;

[0139] FIG3a is a schematic diagram of a possible application framework in a communication system provided in an embodiment of the present application;

[0140] FIG3 b is a schematic diagram of another possible application framework in the communication system provided in an embodiment of the present application;

[0141] FIG4 is a schematic diagram of an encoder and a decoder provided in an embodiment of the present application;

[0142] FIG5 is a schematic diagram of an AI application framework provided in an embodiment of the present application;

[0143] FIG6 is a flow chart of a method for acquiring model data according to an embodiment of the present application;

[0144] FIG7 is a flow chart of another method for acquiring model data provided in an embodiment of the present application;

[0145] FIG8 is a flow chart of another method for acquiring model data provided in an embodiment of the present application;

[0146] FIG9 is a flow chart of another method for acquiring model data provided in an embodiment of the present application;

[0147] FIG10 is a flow chart of another method for acquiring model data provided in an embodiment of the present application;

[0148] FIG11 is a schematic structural diagram of a model data acquisition device provided in an embodiment of the present application;

[0149] FIG12 is a schematic structural diagram of another model data acquisition device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0150] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0151] The present disclosure relates to at least one (item) as follows, indicating one (item) or more (items). More than one (item) refers to two (items) or more than two (items). "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. In addition, it should be understood that although the terms first, second, etc. may be used to describe each object in the present disclosure, these objects should not be limited to these terms. These terms are only used to distinguish each object from each other.

[0152] The terms "including" and "having" and any variations thereof mentioned in the following description of the present disclosure are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes other steps or units that are not listed, or optionally includes other steps or units that are inherent to these processes, methods, products or devices. It should be noted that in the present disclosure, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any method or design described in the present disclosure as "exemplary" or "for example" should not be interpreted as being more preferred or more advantageous than other methods or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.

[0153] The technology provided by the present disclosure can be applied to various communication systems, for example, the communication system can be a fifth generation (5G) or new radio (NR) system, a long term evolution (LTE) system, an LTE frequency division duplex (FDD) system, an LTE time division duplex (TDD) system, a wireless local area network (WLAN) system, a satellite communication system, a future communication system such as a sixth generation (6G) mobile communication system, or a fusion system of multiple systems. The technical solution provided by the present application can also be applied to device to device (D2D) communication, vehicle-to-everything (V2X) communication, machine to machine (M2M) communication, machine type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems.

[0154] A device in a communication system can send a signal to another device or receive a signal from another device. The signal may include information, signaling, or data, etc. The device can also be replaced by an entity, a network entity, a network element, a communication device, a communication module, a node, a communication node, etc. The present disclosure uses the device as an example for description. For example, the communication system may include at least one terminal device and at least one access network device. The access network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the access network device. In addition, it can be understood that if the communication system includes multiple terminal devices, the multiple terminal devices can also send signals to each other, that is, the signal sending device and the signal receiving device can both be terminal devices.

[0155] The information generation method provided in the embodiment of the present application can be applied to wireless communication systems such as 5G, 6G, and satellite communications. Referring to Figure 1, Figure 1 is a simplified schematic diagram of the wireless communication system provided in the embodiment of the present application. As shown in Figure 1, the wireless communication system includes a wireless access network 100. The wireless access network 100 can be a next-generation (e.g., 6G or higher) wireless access network, or a traditional (e.g., 5G, 4G, 3G, or 2G) wireless access network. One or more communication devices (120a-120j, collectively referred to as 120) can be connected to each other or to one or more network devices (110a, 110b, collectively referred to as 110) in the wireless access network 100. Optionally, Figure 1 is only a schematic diagram, and the wireless communication system may also include other devices, such as core network devices, wireless relay devices, and / or wireless backhaul devices, which are not shown in Figure 1.

[0156] Optionally, in actual applications, the wireless communication system may include multiple network devices (also called access network devices) or multiple communication devices at the same time. A network device may serve one or more communication devices at the same time. A communication device may also access one or more network devices at the same time. The embodiments of the present application do not limit the number of communication devices and network devices included in the wireless communication system.

[0157] The network device may be an entity on the network side for transmitting or receiving signals. The network device may be an access device for a communication device to access the wireless communication system in a wireless manner, such as a base station. Base station can broadly cover various names as follows, or be replaced with the following names, such as: NodeB, evolved NodeB (eNB), next generation NodeB (gNB), access network equipment in open radio access network (O-RAN), relay station, access point, transmission point (TRP), transmitting point (TP), master eNodeB (MeNB), secondary eNodeB (SeNB), multi-standard radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit ( The base station may be a macro base station, a micro base station, a relay node, a donor node or the like, or a combination thereof. The network device may also refer to a communication module, a modem or a chip provided in the aforementioned device or apparatus. The network device may also be a mobile switching center and a device to device (Device-to-Device, D2D), vehicle-to-everything (V2X), a device that performs the base station function in machine-to-machine (M2M) communications, a network side device in a 6G network, a device that performs the base station function in a future communication system, etc. The network device may support networks with the same or different access technologies. The embodiments of the present application do not limit the specific technology and specific device form adopted by the network device.

[0158] Network devices can be fixed or mobile. For example, base stations 110a and 110b are stationary and are responsible for wireless transmission and reception in one or more cells from communication device 120. The helicopter or drone 120i shown in Figure 1 can be configured to act as a mobile base station, and one or more cells can move according to the location of the mobile base station 120i. In other examples, the helicopter or drone (120i) can be configured to act as a communication device communicating with base station 110b.

[0159] In the present disclosure, the communication device used to implement the above-mentioned access network functions can be an access network device, a network device that has some of the access network functions, or a device that can support the implementation of the access network functions, such as a chip system, a hardware circuit, a software module, or a hardware circuit and a software module. The device can be installed in the access network device or used in conjunction with the access network device. The method of the present disclosure is described using the example of the communication device used to implement the access network device functions being an access network device.

[0160] A communication device may be an entity on the user side for receiving or transmitting signals, such as a mobile phone. A communication device may be used to connect people, objects, and machines. A communication device may communicate with one or more core networks through a network device. A communication device includes a handheld device with wireless connection capabilities, other processing devices connected to a wireless modem, or an on-board device. A communication device may be a portable, pocket-sized, handheld, computer-built-in, or vehicle-mounted mobile device. The communication device 120 may be widely used in various scenarios, such as cellular communication, device-to-device D2D, vehicle-to-everything V2X, peer-to-peer (P2P), machine-to-machine M2M, machine-type communication MTC, Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, telemedicine, smart grid, smart furniture, smart office, smart wearables, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc. Some examples of the communication device 120 include: user equipment (UE) of the 3GPP standard, fixed devices, mobile devices, handheld devices, wearable devices, cellular phones, smart phones, Session Initialization Protocol (SIP) phones, laptops, personal computers, smart books, vehicles, satellites, Global Positioning System (GPS) devices, target tracking devices, drones, helicopters, aircraft, ships, remote control devices, smart home devices, industrial devices, personal communication service (PCS) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), wireless network cameras, tablet computers, handheld computers, mobile internet devices (MIDs), wearable devices such as smart watches, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, terminals in vehicle networking systems, wireless terminals in self-driving cars, wireless terminals in smart grids, wireless terminals in transportation safety, and smart cities. Wireless terminals in smart cities, such as smart gas pumps, terminal equipment on high-speed trains, and wireless terminals in smart homes, such as smart speakers, smart coffee machines, and smart printers.The communication device 120 can be a wireless device in the above various scenarios or a device used to be set in a wireless device, for example, a communication module, modem or chip in the above devices. The communication device can also be called a terminal, terminal device, user equipment UE, mobile station (MS), mobile terminal (MT), etc. The communication device can also be a communication device in a future wireless communication system. The communication device can be used in a dedicated network device or a general device. The embodiments of the present application do not limit the specific technology and specific device form adopted by the communication device.

[0161] Alternatively, a communication device can function as a base station. For example, a UE can function as a dispatching entity, providing sidelink signals between UEs in V2X, D2D, or P2P scenarios. As shown in Figure 1 , a cell phone 120a and a car 120b communicate with each other using sidelink signals. Cell phone 120a and smart home device 120e communicate without relaying the communication signals through base station 110b.

[0162] In the present disclosure, a communication device for realizing the functions of a communication device may be a terminal device, or a terminal device having some of the functions of the above communication devices, or a device capable of supporting the functions of the above communication devices, such as a chip system, which may be installed in the terminal device or used in combination with the terminal device. In the present disclosure, a chip system may be composed of a chip, or may include a chip and other discrete devices. In the technical solution provided in the present disclosure, the communication device is described as a terminal device or UE as an example.

[0163] Optionally, a wireless communication system is typically composed of cells, with base stations providing cell management and communication services to multiple mobile stations (MS) in the cell. The base station includes a baseband unit (BBU) and a remote radio unit (RRU). The BBU and RRU can be placed in different locations, for example: the RRU is remote and placed in an area with high traffic volume, while the BBU is placed in a central computer room. The BBU and RRU can also be placed in the same computer room. The BBU and RRU can also be different components under the same rack. Optionally, a cell can correspond to a carrier or component carrier.

[0164] It can be understood that the present disclosure can be applied between a network device and a communication device, between a network device and a network device, or between a communication device and a communication device, that is, between a primary device and a secondary device. The primary device can be a network device or a communication device. When the primary device is a network device, the secondary device can be another network device or a communication device. When the primary device is a communication device, the secondary device can be another communication device.

[0165] The following describes the solution using the example of a primary device being a network device, such as an access network device, and a secondary device being a communication device, such as a terminal device. The downlink direction corresponds to the primary device sending data to the secondary device, and the uplink direction corresponds to the secondary device sending data to the primary device.

[0166] Protocol layer structure between access network equipment and terminal equipment

[0167] The communication between the access network device and the terminal device follows a certain protocol layer structure. The protocol layer structure may include a control plane protocol layer structure and a user plane protocol layer structure. For example, the control plane protocol layer structure may include the functions of protocol layers such as the radio resource control (RRC) layer, the packet data convergence protocol (PDCP) layer, the radio link control (RLC) layer, the medium access control (MAC) layer, and the physical layer. For example, the user plane protocol layer structure may include the functions of protocol layers such as the PDCP layer, the RLC layer, the MAC layer, and the physical layer. In one possible implementation, a service data adaptation protocol (SDAP) layer may also be included above the PDCP layer.

[0168] Optionally, the protocol layer structure between the access network device and the terminal may also include an artificial intelligence (AI) layer for transmitting data related to AI functions.

[0169] Taking data transmission between access network equipment and terminal devices as an example, data transmission needs to pass through the user plane protocol layers, such as the SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer. The SDAP layer, PDCP layer, RLC layer, MAC layer, and physical layer can also be collectively referred to as the access layer. Data transmission is divided into sending or receiving based on the direction of transmission, and each of these layers is further divided into a sending part and a receiving part. Taking downlink data transmission as an example, after the PDCP layer obtains data from the upper layer, it transmits the data to the RLC layer and MAC layer. The MAC layer then generates a transport block, which is then wirelessly transmitted through the physical layer. Data is encapsulated accordingly in each layer. For example, data received by a layer from the layer above it is considered a service data unit (SDU) of that layer. After encapsulation by that layer, it becomes a protocol data unit (PDU) and is then passed to the next layer.

[0170] For example, a terminal device may also have an application layer and a non-access layer. The application layer can be used to provide services to applications installed in the terminal device. For example, downlink data received by the terminal device can be sequentially transmitted from the physical layer to the application layer, and then provided by the application layer to the application. For another example, the application layer can obtain data generated by the application and sequentially transmit the data to the physical layer for transmission to other communication devices. The non-access layer can be used to forward user data, such as forwarding uplink data received from the application layer to the SDAP layer, or forwarding downlink data received from the SDAP layer to the application layer.

[0171] Structure of access network equipment

[0172] The access network equipment may include a centralized unit (CU) and a distributed unit (DU). Multiple DUs may be centrally controlled by one CU. As an example, the interface between the CU and the DU may be referred to as an F1 interface. Among them, the control plane (CP) interface may be F1-C, and the user plane (UP) interface may be F1-U. The CU and DU may be divided according to the protocol layers of the wireless network: for example, the functions of the PDCP layer and above protocol layers are set in the CU, and the functions of the protocol layers below the PDCP layer (such as the RLC layer and the MAC layer, etc.) are set in the DU; for another example, the functions of the protocol layers above the PDCP layer are set in the CU, and the functions of the protocol layers below the PDCP layer are set in the DU.

[0173] It is understandable that the above division of the processing functions of CU and DU according to the protocol layer is only an example, and can also be divided in other ways, for example, the CU or DU can be divided into functions with more protocol layers, and for example, the CU or DU can also be divided into partial processing functions with the protocol layer. In one design, some functions of the RLC layer and the functions of the protocol layers above the RLC layer are set in the CU, and the remaining functions of the RLC layer and the functions of the protocol layers below the RLC layer are set in the DU. In another design, the functions of the CU or DU can also be divided according to the service type or other system requirements, for example, by delay, the functions whose processing time needs to meet the delay requirements are set in the DU, and the functions that do not need to meet the delay requirements are set in the CU. In another design, the CU can also have one or more functions of the core network. For example, the CU can be set on the network side to facilitate centralized management. In another design, the RU of the DU is set remotely. Among them, the RU has a radio frequency function.

[0174] Optionally, the DU and the RU may be divided at the physical layer (PHY). For example, the DU may implement high-level functions in the PHY layer, and the RU may implement low-level functions in the PHY layer. When used for transmission, the functions of the PHY layer may include adding cyclic redundancy check (CRC) codes, channel coding, rate matching, scrambling, modulation, layer mapping, precoding, resource mapping, physical antenna mapping, and / or RF transmission functions. When used for reception, the functions of the PHY layer may include CRC, channel decoding, rate matching, descrambling, demodulation, layer mapping, channel detection, resource demapping, physical antenna demapping, and / or RF reception functions. The high-level functions in the PHY layer may include a portion of the functions of the PHY layer, such as a portion of the functions that is closer to the MAC layer, and the low-level functions in the PHY layer may include another portion of the functions of the PHY layer, such as a portion of the functions that is closer to the RF functions. For example, the high-level functions in the PHY layer may include adding CRC codes, channel coding, rate matching, scrambling, modulation, and layer mapping, and the low-level functions in the PHY layer may include precoding, resource mapping, physical antenna mapping, and RF transmission functions; or, the high-level functions in the PHY layer may include adding CRC codes, channel coding, rate matching, scrambling, modulation, layer mapping, and precoding, and the low-level functions in the PHY layer may include resource mapping, physical antenna mapping, and RF transmission functions.

[0175] For example, the functions of the CU can be implemented by one entity, or by different entities. For example, the functions of the CU can be further divided, that is, the control plane and the user plane are separated and implemented by different entities, namely the control plane CU entity (i.e., CU-CP entity) and the user plane CU entity (i.e., CU-UP entity). The CU-CP entity and the CU-UP entity can be coupled with the DU to jointly complete the functions of the access network device.

[0176] In the above architecture, signaling generated by the CU can be sent to the terminal device via the DU, and vice versa. For example, RRC or PDCP layer signaling is ultimately processed into physical layer signaling and sent to the terminal device, or converted from received physical layer signaling. In this architecture, the RRC or PDCP layer signaling can be considered to be sent via the DU, or via the DU and RU.

[0177] Optionally, any of the above-mentioned DU, CU, CU-CP, CU-UP, and RU can be a software module, a hardware structure, or a software module + hardware structure, without limitation. The existence forms of different entities can be different and are not limited. For example, DU, CU, CU-CP, and CU-UP are software modules, and RU is a hardware structure. These modules and their execution methods are also within the scope of protection of this disclosure.

[0178] The access network equipment may support one or more types of fronthaul interfaces, and different fronthaul interfaces correspond to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more baseband functions, and the RU is configured to implement one or more radio frequency functions. If the fronthaul interface between the DU and the RU is another type of interface, relative to the CPRI, some of the downlink and / or uplink baseband functions, such as precoding, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) for downlink, are moved from the DU to the RU for implementation; and for uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / cyclic prefix (CP) removal are moved from the DU to the RU for implementation. In one possible implementation, the interface may be an enhanced common public radio interface (eCPRI). In the eCPRI architecture, the division between the DU and RU is different, corresponding to different types (category, Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, and F.

[0179] Taking eCPRI Cat A as an example, for downlink transmission, based on layer mapping, the DU is configured to implement layer mapping and one or more functions preceding it (i.e., one or more of coding, rate matching, scrambling, modulation, and layer mapping). Other functions after layer mapping (e.g., resource element (RE) mapping, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) addition) are moved to the RU for implementation. For uplink transmission, based on RE demapping, the DU is configured to implement demapping and one or more functions preceding it (i.e., one or more of decoding, rate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, and RE demapping). Other functions after demapping (e.g., one or more of digital BF or fast Fourier transform (FFT) / CP removal) are moved to the RU for implementation. It is understandable that for the functional description of DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol, which will not be described in detail here.

[0180] In one possible design, the processing unit for implementing baseband functions in the BBU is called a baseband high layer (BBH) unit, and the processing unit for implementing baseband functions in the RRU / AAU / RRH is called a baseband low layer (BBL) unit.

[0181] In different systems, CU (or CU-CP and CU-UP), DU or RU may also have different names, but those skilled in the art can understand their meanings. For example, in an open radio access network (open RAN, ORAN) system, CU may also be referred to as O-CU (open CU), DU may also be referred to as O-DU, CU-CP may also be referred to as O-CU-CP, CU-UP may also be referred to as O-CU-UP, and RU may also be referred to as O-RU. Any of the CU (or CU-CP, CU-UP), DU and RU in this application may be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.

[0182] In the embodiments of the present application, the device for implementing the functions of the network device can be a network device; it can also be a device that can support the network device to implement the functions, such as a chip system, a hardware circuit, a software module, or a hardware circuit and a software module. The device can be installed in the network device or used in conjunction with the network device. In the embodiments of the present application, only the device for implementing the functions of the network device is used as an example to illustrate, and does not constitute a limitation on the solutions of the embodiments of the present application.

[0183] The network device and / or terminal device can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; it can also be deployed on the water surface; it can also be deployed on aircraft, balloons and satellites in the air. The embodiments of this application do not limit the scenarios in which the network device and the terminal device are located. In addition, the terminal device and the network device can be hardware devices, or they can be software functions running on dedicated hardware, software functions running on general-purpose hardware, such as virtualization functions instantiated on a platform (e.g., a cloud platform), or entities including dedicated or general-purpose hardware devices and software functions. This application does not limit the specific forms of the terminal device and the network device.

[0184] It should be understood that the number and type of each device in the communication system shown in Figure 1 are for illustration only, and the present disclosure is not limited thereto. In actual applications, the communication system may also include more terminal devices, more access network devices, and other network elements, such as core network devices and / or network elements for implementing artificial intelligence functions. Among them, the network element for implementing artificial intelligence functions may be a RAN intelligent controller (RIC). Figure 2a is a schematic diagram of a communication system applicable to an embodiment of the present application. As shown in Figure 2a, the communication system 200 may include at least one network device, such as the network device 210 shown in Figure 2a; the communication system 200 may also include at least one terminal device, such as the terminal device 220 and the terminal device 230 shown in Figure 2a. The network device 210 and the terminal device (such as the terminal device 220 and the terminal device 230) can communicate via a wireless link. The communication devices in the communication system, such as the network device 210 and the terminal device 220, can communicate using multi-antenna technology.

[0185] Figure 2b is a schematic diagram of another communication system applicable to an embodiment of the present application. Compared to the communication system 200 shown in Figure 2a, the communication system 300 shown in Figure 2b also includes an AI network element 240. AI network element 240 is used to perform AI-related operations, such as constructing a training dataset or training an AI model.

[0186] In one possible implementation, the network device 210 may send data related to the training of the AI ​​model to the AI ​​network element 240, which constructs a training data set and trains the AI ​​model. For example, the data related to the training of the AI ​​model may include data reported by the terminal device. The AI ​​network element 240 may send the results of the operations related to the AI ​​model to the network device 210, and forward them to the terminal device through the network device 210. For example, the results of the operations related to the AI ​​model may include at least one of the following: an AI model that has completed training, an evaluation result or a test result of the model, etc. Exemplarily, a portion of the trained AI model may be deployed on the network device 210, and another portion may be deployed on the terminal device. Alternatively, the trained AI model may be deployed on the network device 210. Alternatively, the trained AI model may be deployed on the terminal device.

[0187] It should be understood that Figure 2b illustrates only the example of a direct connection between AI network element 240 and network device 210. In other scenarios, AI network element 240 may also be connected to a terminal device. Alternatively, AI network element 240 may be connected to both network device 210 and a terminal device simultaneously. Alternatively, AI network element 240 may be connected to network device 210 via a third-party network element (also referred to as a third-party device or third-party entity). This embodiment of the present application does not limit the connection relationship between the AI ​​network element and other network elements.

[0188] The AI ​​network element 240 may also be provided as a module in a network device and / or a terminal device, for example, in the network device 210 or the terminal device shown in FIG. 2 a .

[0189] It should be noted that Figures 2a and 2b are simplified schematic diagrams for ease of understanding. For example, the communication system may also include other devices, such as wireless relay devices and / or wireless backhaul devices, which are not shown in Figures 2a and 2b. In actual applications, the communication system may include multiple network devices and multiple terminal devices. The embodiments of the present application do not limit the number of network devices and terminal devices included in the communication system.

[0190] In order to support AI technology in wireless networks, AI nodes may also be introduced into the network.

[0191] Optionally, the AI ​​node can be deployed in one or more of the following locations in the communication system: access network equipment, terminal equipment, or core network equipment. Alternatively, the AI ​​node can be deployed separately, for example, in a location other than any of the above devices, such as a host or cloud server in an over-the-top (OTT) system. The AI ​​node can communicate with other devices in the communication system, such as one or more of the following: network equipment, terminal equipment, or core network elements.

[0192] It is understood that this application does not limit the number of AI nodes. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on function, such as different AI nodes are responsible for different functions.

[0193] It can also be understood that AI nodes can be independent devices, or they can be integrated into the same device to implement different functions, or they can be network elements in hardware devices, or they can be software functions running on dedicated hardware, or they can be virtualized functions instantiated on a platform (for example, a cloud platform). This application does not limit the specific form of the above-mentioned AI nodes.

[0194] An AI node can be an AI network element or an AI module.

[0195] Figure 3a is a schematic diagram of a possible application framework in a communication system. As shown in Figure 3a, network elements in the communication system are connected through interfaces (such as NG, Xn) or air interfaces. One or more AI modules are provided in one or more devices of these network element nodes, such as core network equipment, access network (radio access network, RAN) nodes, terminals or OAM (for the sake of clarity, only one is shown in Figure 3a). The access network node can be a separate RAN node, or it can include multiple RAN nodes, for example, including CU and DU. The CU and / or DU can also be provided with one or more AI modules. Optionally, the CU can also be split into CU-CP and CU-UP. One or more AI models are provided in the CU-CP and / or CU-UP.

[0196] The AI ​​module is used to implement the corresponding AI function. The AI ​​modules deployed in different network elements may be the same or different. The model of the AI ​​module can implement different functions according to different parameter configurations. The model of the AI ​​module can be configured based on one or more of the following parameters: structural parameters (such as the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of the neuron, the activation function of the neuron, or at least one of the bias in the activation function), input parameters (such as the type of input parameters and / or the dimension of the input parameters), or output parameters (such as the type of output parameters and / or the dimension of the output parameters). Among them, the bias in the activation function can also be called the bias of the neural network.

[0197] An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning, training, or inference processes of different models can be deployed on different nodes or devices, or on the same node or device.

[0198] Figure 3b is a schematic diagram of another possible application framework in a communication system. As shown in Figure 3b, the communication system includes a RAN intelligent controller (RIC). For example, the RIC can be the AI ​​module in Figure 3a, which is used to implement AI-related functions. The RIC includes a near-real-time RIC (near-real time RIC, near-RT RIC) and a non-real-time RIC (non-real time RIC, Non-RT RIC). Among them, the non-real-time RIC mainly processes non-real-time information, such as data that is not sensitive to delay, and the delay of this data can be in the order of seconds. The real-time RIC mainly processes near-real-time information, such as data that is relatively sensitive to delay, and the delay of this data is in the order of tens of milliseconds.

[0199] The near real-time RIC is used for model training and reasoning. For example, it is used to train an AI model and use the AI ​​model for reasoning. The near real-time RIC can obtain network-side and / or terminal-side information from a RAN node (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or a terminal. This information can be used as training data or reasoning data. Optionally, the near real-time RIC can deliver the reasoning result to the RAN node and / or the terminal. Optionally, the reasoning result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the near real-time RIC delivers the reasoning result to the DU, and the DU sends it to the RU.

[0200] The non-real-time RIC is also used for model training and reasoning. For example, it is used to train an AI model and use the model for reasoning. The non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU and / or RU) and / or terminals. This information can be used as training data or reasoning data, and the reasoning results can be submitted to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the non-real-time RIC submits the reasoning results to the DU, and the DU sends it to the RU.

[0201] The near-real-time RIC and non-real-time RIC may also be separately configured as network elements. Optionally, the near-real-time RIC and non-real-time RIC may also be part of other devices. For example, the near-real-time RIC may be configured in a RAN node (e.g., a CU or DU), while the non-real-time RIC may be configured in an OAM, a server (e.g., a cloud server), a core network device, or other network devices.

[0202] It is understandable that all or part of the functions implemented by one or more of the terminal equipment, access network equipment, core network equipment, or network elements for implementing artificial intelligence functions can be virtualized, that is, implemented by one or more of the proprietary processors or general-purpose processors and the corresponding software modules. Among them, since the terminal equipment and the access network equipment involve interfaces for air interface transmission, the transceiver functions of the interfaces can be implemented by hardware. Core network equipment, such as operation administration and maintenance (OAM) network elements, can be virtualized. Optionally, one or more functions of the virtualized terminal equipment, access network equipment, core network equipment, or network elements for implementing artificial intelligence functions can be implemented by cloud devices, such as cloud devices in over the top (OTT) systems.

[0203] The method provided in the present disclosure can be used for communication between access network equipment and terminal equipment, and can also be used for communication between other communication equipment, such as communication between macro base stations and micro base stations in a wireless backhaul link, and communication between two terminal devices in a side link (SL), etc., without limitation.

[0204] To facilitate understanding of the solutions of the embodiments of the present application, the terms that may be involved in the embodiments of the present application are explained below.

[0205] (1) AI model:

[0206] An AI model is an algorithm or computer program that implements AI functionality. It represents the mapping between the model's inputs and outputs. AI models can be neural networks, linear regression models, decision tree models, support vector machines (SVMs), Bayesian networks, Q-learning models, or other machine learning (ML) models.

[0207] (2) Two-end model:

[0208] The two-end model can also be called a bilateral model, collaborative model, dual model, or two-side model. A two-end model is a model composed of multiple sub-models. The sub-models that make up the model must match each other. These sub-models can be deployed on different nodes.

[0209] In one possible design, an embodiment of the present application relates to an encoder for compressing CSI and a decoder for recovering compressed CSI. The encoder and decoder are used in combination, and it can be understood that the encoder and decoder are matching AI models. An encoder can include one or more AI models, and the decoder matched with the encoder also includes one or more AI models. The number of AI models included in the matching encoder and decoder is the same and corresponds one to one.

[0210] In one possible design, a matched set of encoders and decoders can be specifically two parts of the same auto-encoder (AE), as shown in Figure 4. An AE model, in which the encoder and decoder are deployed on different nodes, is a typical bilateral model. The encoder and decoder of an AE model are typically trained together. The encoder processes the input V to produce the processed output z, and the decoder decodes the encoder output z into the desired output V'.

[0211] An autoencoder is a type of neural network that uses unsupervised learning. Its characteristic is that it uses input data as labels, so it can also be understood as a self-supervised learning neural network. Autoencoders can be used for data compression and recovery. For example, the encoder in an autoencoder can compress (encode) data A to obtain data B; the decoder in the autoencoder can decompress (decode) data B to recover data A. Alternatively, the decoder can be understood as the inverse operation of the encoder.

[0212] For example, the AI ​​model in the embodiments of the present application may include an encoder and a decoder. The encoder and decoder are used in combination, and it can be understood that the encoder and decoder are a matching AI model. The encoder and decoder can be deployed on terminal devices and network devices respectively.

[0213] Alternatively, the AI ​​model in the embodiment of the present application may be a single-ended model, which may be deployed on a terminal device or a network device.

[0214] (3) Neural network (NN):

[0215] Neural networks are a specific implementation of AI or machine learning. According to the universal approximation theorem, neural networks can theoretically approximate any continuous function, giving them the ability to learn arbitrary mappings.

[0216] A neural network can be composed of neural units, which can be a computational unit that takes xs and an intercept 1 as input. A neural network is formed by connecting many of these single neural units, meaning that the output of one neural unit can be the input of another. The input of each neural unit can be connected to the local receptive field of the previous layer to extract features from that local receptive field, which can be an area consisting of several neural units.

[0217] Taking the AI ​​model type as a neural network as an example, the AI ​​model involved in this disclosure can be a deep neural network (DNN). Depending on the network construction method, DNN can include feedforward neural networks (FNN), convolutional neural networks (CNN), and recurrent neural networks (RNN).

[0218] (4) Training data set and inference data:

[0219] In the field of machine learning, ground truth usually refers to data that is believed to be accurate or real.

[0220] A training dataset is used to train an AI model. It may include the input to the AI ​​model, or the input and target output of the AI ​​model. A training dataset includes one or more training data. Training data may include training samples input to the AI ​​model, or the target output of the AI ​​model. The target output may also be referred to as a label, sample label, or labeled sample. A label is the true value.

[0221] In the communications field, training datasets can include simulated data collected through simulation platforms, experimental data collected in experimental scenarios, or measured data collected in actual communication networks. Because the geographical environments and channel conditions in which data are generated vary, such as indoor and outdoor locations, mobile speeds, frequency bands, or antenna configurations, the collected data can be categorized during acquisition. For example, data with the same channel propagation environment and antenna configuration can be grouped together.

[0222] Model training essentially involves learning certain characteristics from training data. When training an AI model (such as a neural network), the goal is to ensure that the model's output is as close as possible to the desired predicted value. This is done by comparing the network's predictions with the desired target values. The weight vectors of each layer of the AI ​​model are then updated based on the difference between the two. (Of course, before the first update, there's usually an initialization process, which pre-configures the parameters for each layer of the AI ​​model.) For example, if the network's prediction is too high, the weight vectors are adjusted to predict a lower value. This adjustment is repeated until the AI ​​model predicts the desired target value, or a value very close to it. Therefore, it's necessary to predefine how to compare the difference between the predicted and target values. This is known as the loss function, or objective function. These are important equations used to measure the difference between the predicted and target values. For example, a higher loss function indicates a greater difference. Therefore, training an AI model becomes a process of minimizing this loss, keeping the loss function below a threshold or ensuring that the loss function meets the target requirement. For example, the AI ​​model is a neural network, and adjusting the model parameters of the neural network includes adjusting at least one of the following parameters: the number of layers, width, weights of neurons, or parameters in the activation function of neurons of the neural network.

[0223] Inference data can be used as input to a trained AI model for inference. During the inference process, the inference data is input into the AI ​​model, and the corresponding output is the inference result.

[0224] (5) AI model design:

[0225] The design of an AI model primarily involves data collection (e.g., collecting training data and / or inference data), model training, and model inference. Furthermore, it can also include the application of inference results.

[0226] FIG5 shows an AI application framework.

[0227] In the aforementioned data collection phase, the data source is used to provide training data sets and inference data. In the model training phase, an AI model is obtained by analyzing or training the training data provided by the data source. The AI ​​model represents the mapping relationship between the model's input and output. Learning the AI ​​model through the model training node is equivalent to learning the mapping relationship between the model's input and output using the training data. In the model inference phase, the AI ​​model trained in the model training phase is used to perform inference based on the inference data provided by the data source, obtaining an inference result. This phase can also be understood as inputting the inference data into the AI ​​model and obtaining an output from the AI ​​model, which is the inference result. The inference result can indicate the configuration parameters used (executed) by the execution object and / or the operations performed by the execution object. In the inference result application phase, the inference result is published. For example, the inference result can be centrally planned by the execution (actor) entity, for example, the execution entity can send the inference result to one or more execution objects (e.g., network devices or terminal devices) for execution. Alternatively, the execution entity can provide feedback on the model's performance to the data source to facilitate subsequent model update and training.

[0228] It is understandable that a communication system may include network elements with artificial intelligence capabilities. The above-mentioned AI model design-related steps can be performed by one or more network elements with artificial intelligence capabilities. In one possible design, AI functions (such as AI modules or AI entities) can be configured in existing network elements in the communication system to implement AI-related operations, such as AI model training and / or inference. For example, the existing network element can be a network device or a terminal device. Alternatively, in another possible design, an independent network element can be introduced into the communication system to perform AI-related operations, such as training an AI model. The independent network element can be referred to as an AI network element, an AI node, or an AI entity, etc., and the embodiments of the present application are not limited to these names. For example, the AI ​​network element can be directly connected to the network equipment in the communication system, or it can be indirectly connected to the network equipment through a third-party network element. The third-party network element can be a core network element such as an authentication management function (AMF) network element, a user plane function (UPF) network element, an operation administration and maintenance (OAM) network element, a server (such as a cloud server), an over-the-top (OTT) device, or other network element, without limitation. Exemplarily, the independent AI network element, AI entity, or AI node can be deployed on one or more of the network device side, the terminal device side, or the core network side. Optionally, it can be deployed on a server, such as a cloud server, or an OTT device, or other device. Exemplarily, an AI network element 240 is introduced into the communication system shown in FIG2b. It can be understood that the aforementioned AI module, AI entity, AI network element, or AI node can be used to perform one or more of the AI ​​functions, where the AI ​​functions may include: processing of AI models, such as training and / or updating of AI models, monitoring of AI models, management of AI models, such as registration and / or deregistration of AI models, or application reasoning of AI models.

[0229] The training process of different models can be deployed in different devices or nodes, or in the same device or node. The inference process of different models can be deployed in different devices or nodes, or in the same device or node. Taking the completion of the model training phase of a terminal device as an example, the terminal device can train the matching encoder and decoder, and then send the model parameters of the decoder to the network device. Taking the completion of the model training phase of a network device as an example, after the network device trains the matching encoder and decoder, it can indicate the model parameters of the encoder to the terminal device. Taking the completion of the model training phase of an independent AI network element as an example, the AI ​​network element can train the matching encoder and decoder, and then send the model parameters of the encoder to the terminal device and the model parameters of the decoder to the network device. Then, the model inference phase corresponding to the encoder is performed in the terminal device, and the model inference phase corresponding to the decoder is performed in the network device.

[0230] Among them, the model parameters may include one or more of the following structural parameters of the model (such as the number of layers and / or weights of the model, etc.), the input parameters of the model (such as input dimension, number of input ports), or the output parameters of the model (such as output dimension, number of output ports). It can be understood that the input dimension may refer to the size of an input data. For example, when the input data is a sequence, the input dimension corresponding to the sequence may indicate the length of the sequence. The number of input ports may refer to the number of input data. Similarly, the output dimension may refer to the size of an output data. For example, when the output data is a sequence, the output dimension corresponding to the sequence may indicate the length of the sequence. The number of output ports may refer to the number of output data.

[0231] (6) Channel information:

[0232] In a communication system (for example, an LTE communication system or an NR communication system, etc.), a network device needs to determine the resource, MCS, and precoding configuration of the downlink data channel of the scheduling terminal device based on channel information. It can be understood that channel information can also be referred to as channel state information (CSI) or channel environment information, which is information that can reflect channel characteristics and channel quality. In this application, the meaning of CSI is broader than that of CSI in traditional schemes, and is not limited to channel quality indication (CQI), precoding matrix indicator (PMI), rank indicator (RI), or CSI-RS resource indicator (CRI). It can also be channel response information (such as a channel response matrix), weight information corresponding to the channel response, reference signal receiving power (RSRP), or signal to interference plus noise ratio (SINR). One or more of the above.

[0233] CSI measurement refers to the receiver solving the channel information based on the reference signal sent by the transmitter, that is, estimating the channel information using the channel estimation method. Exemplarily, the reference signal may include one or more of a channel state information reference signal (CSI-RS), a synchronization signal / physical broadcast channel block (SSB), a sounding reference signal (SRS), or a demodulation reference signal (DMRS). CSI-RS, SSB, and DMRS can be used to measure downlink CSI. SRS and DMRS can be used to measure uplink CSI.

[0234] Taking FDD communication scenarios as an example, in FDD communication scenarios, because uplink and downlink channels are not reciprocal or cannot be guaranteed, network equipment typically transmits a downlink reference signal to the terminal device. The terminal device performs channel and interference measurements based on the received downlink reference signal to estimate the downlink CSI. The terminal device generates a CSI report based on a protocol predefined method or a network device configuration method and feeds it back to the network device to obtain the downlink CSI.

[0235] Exemplarily, the CSI may include at least one of the following: channel quality indication (CQI), precoding matrix indicator (PMI), rank indicator (RI), CSI-RS resource indicator (CRI), layer indicator (LI), reference signal receiving power (RSRP), or signal to interference plus noise ratio (SINR). The signal to interference plus noise ratio may also be referred to as the signal to interference plus noise ratio.

[0236] The RI indicates the number of downlink transmission layers recommended by the terminal device, the CQI indicates the modulation and coding scheme supported by the current channel conditions as determined by the terminal device, and the PMI indicates the precoding recommended by the terminal device. The number of precoding layers indicated by the PMI corresponds to the RI.

[0237] It should be understood that the RI, CQI, and PMI indicated in the above CSI report are only recommended values ​​for the terminal device, and the network device may perform downlink transmission according to part or all of the information indicated in the CSI report. Alternatively, the network device may not perform downlink transmission according to the information indicated in the CSI report.

[0238] The introduction of AI technology into wireless communication networks has resulted in a CSI feedback method based on AI models. Terminal devices use AI models to compress and feedback CSI, and network equipment uses AI models to recover the compressed CSI. AI-based CSI feedback transmits a sequence (such as a bit sequence), resulting in lower overhead than traditional CSI feedback.

[0239] Taking Figure 4 as an example, the encoder in Figure 4 can be a CSI generator, and the decoder can be a CSI reconstructor. The encoder can be deployed in a terminal device, and the decoder can be deployed in a network device. The terminal device can use the encoder to generate CSI feedback information z from the original CSI information V. The terminal device then reports a CSI report, which can include the CSI feedback information z. The network device can use the decoder to reconstruct the CSI information, thereby obtaining the recovered CSI information V'.

[0240] The CSI original information V may be obtained by the terminal device through CSI measurement. For example, the CSI original information V may include the channel response of the downlink channel or the eigenvector matrix (a matrix composed of eigenvectors) of the downlink channel. The encoder processes the eigenvector matrix of the downlink channel to obtain CSI feedback information z. In other words, the compression and / or quantization operation of the eigenmatrix according to the codebook in the related scheme is replaced by the operation of processing the eigenmatrix by the encoder to obtain CSI feedback information z. The terminal device reports the CSI feedback information z. The network device processes the CSI feedback information z through the decoder to obtain CSI recovery information V'.

[0241] The following further illustrates the training process and reasoning process of the AI ​​model in the embodiments of the present application.

[0242] The training data used to train AI models includes training samples and sample labels. For example, the training samples are channel information determined by the terminal device, and the sample labels are the actual channel information, i.e., the true value CSI. If the encoder and decoder belong to the same autoencoder, the training data can only include the training samples, or the training samples are the sample labels.

[0243] In the field of wireless communications, the true CSI may be high-precision CSI.

[0244] The specific training process is as follows: the model training node uses an encoder to process channel information, that is, training samples, to obtain channel feedback information, such as CSI feedback information, and uses a decoder to process the feedback information to obtain recovered channel information, that is, channel recovery information, such as CSI recovery information. Then, the difference between the channel recovery information and the corresponding sample label is calculated, that is, the value of the loss function, and the parameters of the encoder and decoder are updated according to the value of the loss function, so that the difference between the recovered channel information and the corresponding sample label is minimized, that is, the loss function is minimized. Exemplarily, the loss function can be the minimum mean square error (MSE) or cosine similarity. Repeat the above operations to obtain an encoder and decoder that meet the target requirements. The above model training node can be a terminal device, a network device, or other network elements with AI functions in a communication system.

[0245] It should be understood that the above description uses the AI ​​model for CSI compression as an example. The AI ​​model can also be used in other scenarios in CSI feedback. For example, the AI ​​model can be used for CSI prediction, that is, predicting channel information at one or more future moments based on channel information measured at one or more historical moments. The embodiments of this application do not limit the specific use of the AI ​​model in CSI feedback scenarios.

[0246] It should be understood that, in this application, indication includes direct indication (also known as explicit indication) and implicit indication. Direct indication of information A refers to including information A; implicit indication of information A refers to indicating information A through the correspondence between information A and information B and the direct indication of information B. The correspondence between information A and information B can be predefined, pre-stored, pre-burned, or pre-configured.

[0247] It should be understood that, in this application, information C is used to determine information D, which includes both information D being determined solely based on information C and information D being determined based on information C and other information. Furthermore, information C can also be used to determine information D indirectly, for example, where information D is determined based on information E, and information E is determined based on information C.

[0248] In addition, in each embodiment of the present application, "network element A sends information A to network element B" can be understood as the destination end of the information A or the intermediate network element in the transmission path between the destination end and the network element B, which may include directly or indirectly sending information to network element B. "Network element B receives information A from network element A" can be understood as the source end of the information A or the intermediate network element in the transmission path between the source end and the network element A, which may include directly or indirectly receiving information from network element A. The information may be processed as necessary between the source end and the destination end of the information transmission, such as format changes, but the destination end can understand the valid information from the source end. Similar expressions in this application can be understood similarly and will not be elaborated here.

[0249] The method of the embodiment of the present application is described in detail below.

[0250] 6 , which is a flow chart of a model data acquisition method provided by an embodiment of the present application. Optionally, the method can be applied to the aforementioned model data acquisition system, such as the model data acquisition system shown in FIG1 . It can be understood that the example shown in FIG6 is an example of the execution subject of the interaction diagram using a first node (data synthesis node, or it can be called a central node, a first entity, etc.) and a second node (data providing node, or it can be called a node to be served, a second entity, etc.) as an example. It can further include model processing (such as model training or updating, etc.) nodes and model use (such as model reasoning, etc.) nodes. Among them, the model processing node can be the first node, or it can be other nodes. The model use node can be the second node, or it can be other nodes, etc. The model data acquisition method shown in FIG6 may include steps 601-602. Steps 601-602 are as follows:

[0251] 601. A second node sends first information to a first node, where the first information includes the following types of information: original data and characteristic information of the original data. Correspondingly, the first node receives the first information.

[0252] Exemplarily, the first node may be a data synthesis node, or may be referred to as a data generation node, a central node, a first entity, etc. The second node may be a data provision node, or may be referred to as a node to be served, a second entity, etc. Optionally, the first node is an access network device, and the second node is a terminal device.

[0253] This raw data refers to uncompressed or uninformative measured data. By providing the raw data to the first node, the second node helps ensure that the synthetic data generated by the first node has the same dimensions and accuracy as the original data. Synthetic data is non-artificially created data that mimics real-world data. It is created using computational algorithms and simulations based on generative artificial intelligence techniques. Synthetic datasets share the same mathematical properties as the actual data on which they are based, but do not contain the same information.

[0254] The feature information of the original data is the information obtained by extracting features from the original data. This feature information extracts key information from the original data. By providing the feature information of the original data to the first node, the second node enables the generative model to understand which features the generated data should be similar to the original data.

[0255] The characteristic information of the original data may include characteristic quantity information. Alternatively, the characteristic information of the original data may include not only characteristic quantity information but also characteristic information corresponding to the characteristic quantity. For example, for channel information, the characteristic quantity information may include the channel angle distribution, the channel delay distribution, etc.; the characteristic information corresponding to the characteristic quantity may include the value of the channel angle distribution, the value of the channel delay distribution, etc.

[0256] The type of the first information refers to the type of information included in the first information. For example, if the type of the first information is A, the first information includes the original data and the characteristic information of the original data; if the type of the first information is B, the first information includes the original data; and if the type of the first information is C, the first information includes the characteristic information of the original data.

[0257] Optionally, the first information further includes at least one of the following types of information: the priority of the original data, and the data label of the original data.

[0258] By reporting the priority of the original data, the first node determines the number of similar synthetic data to be generated based on the priority or probability of each original data, so that the distribution of the synthetic data is consistent with expectations, for example, the distribution of the synthetic data is consistent with the scenario in which the original data is applied. For example, taking channel information as an example, if the second node is in an outdoor environment and the obstruction from the base station to the user equipment is less than that indoors, the measured channel information based on the line-of-sight (LOS) path of the wireless signal can be given a higher priority, while the measured channel based on the non-line-of-sight (NLOS) path of the wireless signal can be given a lower priority, so that the synthesized channel information can be based on the sparse channel under the LOS path.

[0259] The data label of the raw data can be understood as the category to which the raw data belongs, obtained by classifying the raw data. For example, if the raw data is classified according to the scene to which it belongs, the label is the scene classification label.

[0260] Data labeling (or data annotation) is part of the preprocessing phase when developing machine learning (ML) models. It is responsible for identifying raw data (such as images, text files, and videos). One or more data labels can then be added to the raw data to specify the model's context and help the machine learning model make accurate predictions. Data labeling supports a variety of different machine learning and deep learning use cases, including computer vision and natural language processing (NLP).

[0261] The priority of the original data can be taken from the predefined P file. For example, the meaning of each priority file can be predefined. For example, for the original data with priority file P1, the generation model generates N1 similar synthetic data, and for the original data with priority file P2, the generation model generates N2 similar synthetic data. Alternatively, the number of synthetic data generated by the generation model that are similar to the original data with priority file P1 is x1, and the number of synthetic data generated by the generation model that are similar to the original data with priority file P2 is x2, and the ratio of x1 to x2 is equal to the ratio of N1 to N2, where N1 and N2 are predefined values.

[0262] Exemplarily, the second node sends first information to the first node, where the first information includes: original data, characteristic information of the original data, and a data label of the original data. For example, taking the original data as channel state information, the characteristic information of the original data may be an extracted channel characteristic vector (one or more characteristic vectors obtained after performing singular value decomposition (SVD) on the original channel), or compressed information after compressing the channel characteristic vector. The data label of the original data may be channel quality information (Channel Quality Indication (CQI), Rank Indication (RI), etc.), or scenario information (channel NLOS / LOS classification information, user channel high / medium / low speed classification information, or other scenario classification information), etc.

[0263] The second node sends the first information to the first node so that the first node generates multiple synthetic data based on the first information. The second node uses the characteristic information of the original data as a condition for synthesizing the data, so that the synthetic data imitates the original data in some specific characteristic dimensions, which can improve the diversity of the synthetic data and also improve the accuracy of the synthetic data.

[0264] In a possible implementation, step 603 may also be included, in which the first node sends second information to the second node. The second information indicates a manner in which the second node obtains the first information, such as an acquisition path, an acquisition source, or one or more of parameters involved in the acquisition path or source. For example, when the first information includes characteristic information of the original data, the manner in which the second node obtains the first information includes a manner in which the characteristic information of the original data is acquired. That is, the first node indicates to the second node a manner in which the characteristic information of the original data is acquired, so that the second node processes the original data based on the manner in which the characteristic information of the original data is acquired to obtain the characteristic information of the original data.

[0265] For example, the method for obtaining the characteristic information of the original data may be D1: extracting N characteristic vectors of the original data, or D2: extracting N characteristic vectors from the characteristic vectors corresponding to the first M largest eigenvalues ​​of the original data matrix. For another example, D3: the method for obtaining the characteristic information of the original data is compressed information obtained by using the R16 / R17 codebook method. For another example, D4: the method for obtaining the characteristic information of the original data is compressed information obtained by using an AI-based encoder. For another example, the method for obtaining the characteristic information of the original data may also be a feature quantity to instruct the second node to calculate the feature information corresponding to the corresponding feature quantity. Specifically, the second information may include an acquisition method identifier (for example, D1 / D2 / D3 / D4), or may include parameters under a certain acquisition method (for example, for method D1, it may indicate the number of characteristic vectors N; for method D4, it may indicate the identifier of the AI ​​encoder).

[0266] Optionally, when the first information includes the raw data, the method by which the second node obtains the first information includes a method for acquiring the raw data. For example, the method for acquiring the raw data may be acquiring specific data within a specified time period, such as from 9:00 to 15:00 on a particular day, and the specific data may be, for example, temperature data, humidity data, or text data. For another example, the method for acquiring the raw data may be acquiring raw data from a preset area using a preset sensor. The preset sensor may be, for example, a camera, a lidar, or the like.

[0267] Optionally, when the first information includes the data label of the original data, the way in which the second node obtains the first information also includes a way to obtain the data label of the original data. For example, the second information indicates the way or standard by which the second node determines the data label of the data. Taking channel information as an example, E1: indicates that the data label is determined according to sparsity (for example, according to a threshold value of the sparsity of the channel, the channel label is determined to be LOS or NLOS); or, E2: determines the data label according to the delay distribution, etc.; E3: or directly indicates the label, for example, indicates that the data label is the CQI value of the channel, E4: or indicates that the data label is the RI of the channel, etc. Specifically, the second information may include an acquisition method identifier (for example, E1 / E2 / E3 / E4).

[0268] The first node sends the second information to the second node, so that the second node processes the original data based on the second information to obtain the first information.

[0269] In a possible implementation, step 604 may also be included, where the first node further sends third information to the second node, and the third information indicates a reporting configuration of the first information. Exemplarily, the reporting configuration includes at least one of the following: the type of the first information, the respective quantity or total quantity of one or more types of the first information. That is, the first node sends the third information to the second node so that the second node reports based on the type of the first information and / or the quantity of the first information indicated in the third information. For example, multiple types of first information are predefined. Furthermore, when the third information indicates that the first information is a type identifier, the second node can determine the type of information included in the first information. For example, according to the type identifier of the first information predefined above, the third information indicates that the type of the first information is A, then it can be known that the first information includes the above-mentioned original data and the characteristic information of the original data; or, the information type identifier contained in the first information is predefined, and the identifier of the original data type is defined as A_1, the identifier of the original data characteristic information type is defined as A_2, the identifier of the tag type is defined as A_3, and the identifier of the priority type is defined as A_4, then the third information indicates the identifier of the information type contained in the first information, for example, the third information indicates that the type of the first information is {A_1, A_2, A_3}, then the second node can determine that the first information includes the above-mentioned original data and the characteristic information of the original data. information and labels of the original data; for another example, when the third information also indicates that the data volume of the first information is 1000, it means that the second node needs to report the data volume of the above-mentioned original data and the characteristic information of the original data as 1000 respectively; or it can also indicate that the sum of the data volume of the original data and the characteristic information of the original data is 1000, etc., wherein, for example, the second node reports the data volume of the above-mentioned original data as 500 and the data volume of the characteristic information of the original data as 500, or the data volume of the characteristic information of the original data is more or less than the data volume of the original data, etc., such as the ratio of the two is 6:4, and the ratio can be predefined or configured, and this solution does not limit this. For another example, multiple values ​​of a quantity can be predefined, and the third information enables the second node to determine the respective quantity or total quantity of one or more types of first information by indicating the identifier of the predefined quantity value.

[0270] Optionally, the reporting configuration may further include one or more items of the time domain resource information or the frequency domain resource information of the first information, which is not limited in this solution.

[0271] The first node sends the third information to the second node, so that the second node reports the first information based on the reporting configuration of the first information indicated in the third information. This allows the separate data provider and the data synthesizer to align their understanding of the data, thereby improving the quality of the synthesized data.

[0272] In one possible implementation, at least two of the first information, the second information, and the third information may be different contents included in the same information sent by the first node, or the first information, the second information, and the third information may be contents of different messages sent by the first node. This solution is not limited to this.

[0273] It can be understood that the content not indicated in the above first information, second information and third information, such as reported time domain resource information or frequency domain resource information, etc., can be predefined by the protocol, or pre-set or stored in the first node and the second node, such as a default configuration, or determined based on other information.

[0274] Based on the indication of the above-mentioned second information and / or third information, the second node processes the original data and then sends the first information to the first node. For example, the third information indicates that the type of the first information is A, that is, the first information includes the above-mentioned original data and the characteristic information of the original data; the second information indicates that the method for obtaining the characteristic information of the original data is to extract N characteristic vectors of the original data, then the second node determines the original data and the characteristic information of the original data, and reports it according to the indication of the third information. Optionally, if the third information indicates that the first information contains the characteristic information of the original data, but the second node does not receive an indication of the method for obtaining the characteristic information, a predefined method for obtaining the characteristic information can be used. Similarly, if the third information indicates that the first information contains the data label of the original data, but the second node does not receive an indication of the method for obtaining the data label of the original data, a predefined method for obtaining can be used.

[0275] In one possible implementation, step 605 may further include the first node sending fourth information to the second node, where the fourth information indicates that the first node has the capability to jointly process feature information of multiple raw data in the same group. The joint processing may include, for example, weighted summing of the feature information of multiple raw data in the same group.

[0276] The first node sends the fourth information to the second node so that the second node can divide the original data that can be jointly processed into the same group, while dividing the original data that cannot be jointly processed into the same group, thereby avoiding the fusion of original data from different groups to generate synthetic data that does not meet expectations.

[0277] Optionally, the first information further includes the following types of information: group information of the original data. The second node reports the group information of the original data to the first node, and then the first node can jointly process feature information of multiple original data in the same group.

[0278] It is understood that in this example, the raw data that can be jointly processed are divided into the same group. This group information may be different from the data labels of the raw data. For example, the criteria for dividing the raw data into groups and the criteria for dividing the raw data data labels may be different. Of course, the criteria for dividing the raw data into groups may also be the same, that is, the group information is the data label, and this solution does not impose any restrictions on this.

[0279] It can be understood that step 603, step 604 and step 605 in FIG6 are one or more optional steps. The steps shown in FIG6 are only examples, and there is no limitation on the order of execution and the time of execution.

[0280] 602. The first node generates a plurality of synthetic data based on the first information, and the plurality of synthetic data are used as inputs of the model.

[0281] The first node generates multiple synthetic data based on the original data and the characteristic information of the original data. For example, the first node inputs the original data and the characteristic information of the original data into a generation model to generate multiple synthetic data. Alternatively, the first node preprocesses the first information and then inputs the processed information into the generation model to generate multiple synthetic data. The preprocessing may, for example, normalize the first information or concatenate the original data and the characteristic information of the original data in the first information.

[0282] Among them, in the traditional data generation method, the original measured data is used as the condition for synthesizing data, and the generation model will imitate the original data, including aspects such as data dimension accuracy and content, so that the synthesized data is highly similar to the measured data. However, using only the original data as input can easily lead to insufficient diversity of the synthesized data. In the solution of the present application, the feature information of the original data, or the feature information of the original data and the data label are used as the conditions for synthesizing data, and the generation model imitates the original data in some specific feature dimensions, which can improve the diversity of the synthesized data. Therefore, further, based on the joint representation of the original data and the feature information of the original data, or the joint representation of the original data, the feature information of the original data and the data label of the original data as the conditions for synthesizing data, data generation is performed with a small amount of data, which can simultaneously meet the requirements of high diversity of synthesized data and close feature distribution of synthesized data to that of the original data.

[0283] For example, taking channel generation as an example, the original data is the channel information, and the characteristic information of the original data may include the channel angle power spectrum characteristics, the channel delay spectrum characteristics, etc. Then the generation model can know that the generated channel should be similar to the original channel information in the angle and delay power spectrum distribution, while other characteristic values, such as the randomness of Doppler frequency deviation, can be higher.

[0284] Another example is weighted summing of the characteristic information of multiple raw data from the same group to increase the diversity of the synthesized data. For example, taking channel generation as an example, the raw data is channel information, and the characteristic information of the raw data may include channel angular power spectrum characteristics, channel delay spectrum characteristics, etc. By weighted summing the characteristic information of multiple channels with different angular power spectrum distributions and inputting it into the generation model, channels with more diverse angular power spectrum distributions can be synthesized, thereby further increasing the diversity and accuracy of the synthesized data.

[0285] Among them, the above-mentioned multiple synthetic data can be used as input of the model, and the input can be used for model processing, such as model training, model updating, model reasoning or model performance monitoring.

[0286] Model training, updating, and other processing can be performed on the first node, or the synthesized data can be transmitted to a model processing node for model training, updating, and other processing. After model training or updating is completed, model inference can be performed, or the model can be transmitted to a model-using node for model inference. For example, it can be transmitted to a second node for model inference.

[0287] In a possible implementation manner, the first node further sends at least a subset of the plurality of synthesized data to the second node.

[0288] At least one subset of the multiple synthetic data can be, for example, a subset of the synthetic data obtained by randomly sampling multiple samples from the multiple synthetic data. For another example, multiple samples are randomly sampled from the synthetic data obtained based on at least one of the multiple data labels to obtain at least one subset of the synthetic data. One possible implementation is that the data label of the original data has three possible values ​​(for example, data label a, data label b and data label c), one subset of the synthetic data is a subset obtained by randomly sampling multiple samples from the synthetic data obtained by inputting data label a as a condition, another subset of the synthetic data is a subset obtained by randomly sampling multiple samples from the synthetic data obtained by inputting data label b (or data label c) as a condition, and so on. Alternatively, a subset of the synthetic data can be obtained by randomly sampling multiple samples S1 from the synthetic data obtained with data label a as the input condition, randomly sampling multiple samples S2 from the synthetic data obtained with data label b as the input condition, and randomly sampling multiple samples S3 from the synthetic data obtained with data label c as the input condition. Then, based on the multiple samples S1, multiple samples S2, and multiple samples S3, one subset is obtained, and this method is repeated to obtain another subset. Of course, the subset can also be obtained by other methods, and this solution does not limit this.

[0289] Based on the received at least one subset, the second node sends fifth information to the first node. The fifth information indicates an evaluation result of at least one synthetic data item in the at least one subset. The evaluation result indicates at least one synthetic data item in the at least one subset that needs to be removed, or indicates at least one synthetic data item in the at least one subset that does not need to be removed. Accordingly, the first node receives the fifth information.

[0290] The synthetic data that needs to be eliminated can, for example, be synthetic data that has a different feature distribution from the original data of the second node. In one possible implementation, the second node obtains the evaluation result by determining whether the received synthetic data is similar to its actual measured data. For example, the second node calculates the distance between a synthetic data in at least one subset and its local data. If there is a local data such that the distance between the synthetic data and the local data satisfies a first condition, the synthetic data is judged to have the same feature distribution as the original data and is not eliminated; if the distance between the synthetic data and all local data does not meet the first condition, the synthetic data is judged to have a different feature distribution from the original data and needs to be eliminated. The first condition can be a specific range or a threshold value. Optionally, the distance between the synthetic data and the local data can be calculated by taking the norm of the difference between the synthetic data and the local data as the distance between the synthetic data and the local data. Alternatively, the data feature information of the synthetic data and the local data is extracted separately, and the norm of the difference between the data feature information of the synthetic data and the local data is taken as the distance between the synthetic data and the local data.

[0291] Optionally, the fifth information may be 0 or 1 to indicate whether each synthesized data is to be removed. Alternatively, the second node may use a list of synthesized data identifiers to indicate the synthesized data to be removed from the synthesized data subset, or use a list of synthesized data identifiers to indicate the synthesized data not to be removed from the synthesized data subset. Each synthesized data in the synthesized data subset has a unique identifier.

[0292] The second node may also obtain the above evaluation result based on other methods, which is not limited in this solution.

[0293] The first node filters the plurality of synthesized data based on the received evaluation result, thereby obtaining at least one filtered synthesized data.

[0294] Optionally, the at least one synthesized data after screening is determined based on a distance between the first synthesized data to be eliminated and other synthesized data in the plurality of synthesized data. The first synthesized data to be eliminated is the synthesized data to be eliminated determined according to the fifth information.

[0295] For example, the first node calculates the distance between the synthetic data to be removed and other synthetic data in the first node. If there is a synthetic data such that the distance between the synthetic data and the synthetic data to be removed is not greater than a certain threshold, then the synthetic data is determined to be removed as well.

[0296] Screening synthetic data through the synthetic data evaluation process can effectively improve the quality of synthetic data.

[0297] The at least one synthesized data after screening is used as the input of the model. For the introduction of this part, please refer to the description of step 601, which will not be repeated here.

[0298] Optionally, the first node is a third-party device that performs the aforementioned actions related to the first node. For example, the above steps 601 and 602 are both performed by the third-party device.

[0299] Optionally, the second node is a third-party device that performs the aforementioned second-node-related actions. For example, the above step 601 is performed by the third-party device.

[0300] Optionally, the first node is a network device. For example, steps 601 and 602 are both performed by the network device.

[0301] Optionally, the second node is a terminal device. For example, the above step 601 is performed by the terminal device.

[0302] Optionally, the first node includes a network device and a third-party device. In one example, step 602 can be performed by a third-party device, such as an OTT or cloud server, and step 601 can be performed by the network device. Furthermore, the network device and the third-party device can also communicate with each other to transmit the content transmitted in step 601.

[0303] In an embodiment of the present application, a first node receives raw data and feature information of the raw data from a second node; then, based on the raw data and feature information, generates multiple synthetic data sets. This multiple synthetic data sets are used for model processing. In this example, synthetic data that meets diversity requirements and has the same data distribution as the measured data can be generated based on the raw data and feature information of the raw data, thereby improving the quality of the synthetic data and thereby helping to improve the performance of model training or updating.

[0304] The above example is introduced by taking the example of a first node sending a second information to a second node, where the second information indicates a method for obtaining characteristic information of the original data. Alternatively, the second information can also be sent by a third node to the second node (the embodiment shown in Figure 7). That is, the third node indicates to the second node a method for obtaining characteristic information of the original data, so that the second node processes the method based on the method for obtaining characteristic information of the original data to obtain characteristic information of the original data. Among them, the third node can be a model training node or a model use node, and can also be called a third entity, such as a model training entity or a model use entity.

[0305] The above description uses the example of the second node processing the raw data to obtain the characteristic information of the raw data. Alternatively, the first node may process the raw data to obtain the characteristic information of the raw data (the embodiment shown in FIG8 and FIG9 ), or the third node may process the raw data to obtain the characteristic information of the raw data (the embodiment shown in FIG10 ).

[0306] The above alternative solutions are described in detail below based on Figures 7, 8, 9 and 10 respectively.

[0307] With reference to Figure 7, it is a flow chart of another model data acquisition method provided by an embodiment of the present application. The example shown in Figure 7 is illustrated by taking the first node (data synthesis node, or it can be called a central node, a first entity, etc.), the second node (data providing node, or it can be called a node to be served, a second entity, etc.) and the third node (model training node or model using node, or it can be called a third entity, such as a model training entity or a model using entity) as an example of the execution subject of the interaction diagram. The difference between this example and the example shown in Figure 6 is that the second information is sent by the third node to the second node, rather than by the first node to the second node. The model data acquisition method shown in Figure 7 may include steps 701-705. Steps 701-705 are specifically as follows:

[0308] 701. A first node sends third information to a second node, where the third information indicates a reporting configuration of the first information. Correspondingly, the second node receives the third information.

[0309] For the introduction of this part, please refer to the description of step 601 in the embodiment shown in FIG6 , which will not be repeated here.

[0310] 702. The third node sends second information to the second node, where the second information indicates a method for obtaining characteristic information of the original data. Correspondingly, the second node receives the second information.

[0311] The third node determines a feature extraction method for the original data according to the model requirements, and then sends the method for obtaining the feature information of the original data to the second node.

[0312] 703. The second node processes the original data based on the second information to obtain feature information of the original data.

[0313] For the introduction of this part, please refer to the description of step 601 in the embodiment shown in FIG6 , which will not be repeated here.

[0314] Optionally, in a first possible implementation, the third node further sends a method for obtaining the data label of the original data to the second node. For example, the second information further indicates a method for obtaining the data label of the original data.

[0315] In a second possible implementation, the first node sends second information to the second node, where the second information indicates how the second node processes the first information, and the processing method for the second node to process the first information includes a method for obtaining a data tag of the original data.

[0316] 704. The second node sends first information to the first node, where the first information includes the following types of information: original data and characteristic information of the original data. Correspondingly, the first node receives the first information.

[0317] Optionally, the first information also includes the following types of information: data labels of original data.

[0318] For the introduction of this part, please refer to the description of step 601 in the embodiment shown in FIG6 , which will not be repeated here.

[0319] 705. The first node generates a plurality of synthetic data based on the first information, and the plurality of synthetic data are used for processing the model.

[0320] For the introduction of this part, please refer to the description of step 602 in the embodiment shown in FIG. 6 , which will not be repeated here.

[0321] Optionally, the first node is a third-party device that performs the aforementioned actions related to the first node. For example, steps 701, 704, and 705 are all performed by the third-party device.

[0322] Optionally, the second node is a third-party device that performs the aforementioned second-node-related actions. For example, steps 701, 702, 703, and 704 are performed by a third-party device.

[0323] Optionally, the third node is a third-party device that performs the aforementioned third-node-related actions. For example, the above step 702 is performed by the third-party device.

[0324] Optionally, the first node is a network device. For example, steps 701, 704, and 705 are all performed by the network device.

[0325] Optionally, the second node is a terminal device. For example, the above steps 701, 702, 703, and 704 are performed by the terminal device.

[0326] Optionally, the third node is an AI node. For example, the above step 702 is performed by the AI ​​node.

[0327] Optionally, the first node includes a network device and a third-party device. In one example, step 705 can be performed by a third-party device, such as an OTT or cloud server, and steps 701 and 703 can be performed by the network device. Furthermore, the network device and the third-party device can communicate with each other to transmit the content transmitted in steps 701 and 703.

[0328] Optionally, the second node includes a terminal device and a third-party device. In one example, step 703 can be performed by a third-party device, such as an OTT or cloud server, and steps 701, 702, and 704 can be performed by the terminal device. Furthermore, the terminal device and the third-party device can communicate with each other to transmit the content transmitted in steps 701, 702, and 704.

[0329] Optionally, the third node is an AI entity. In one example, the above step 702 can be performed by a third-party device, such as an OTT or a cloud server.

[0330] In an embodiment of the present application, the second node processes the raw data based on the method for obtaining the characteristic information of the raw data from the third node to obtain the characteristic information of the raw data. Furthermore, the second node sends the raw data and the characteristic information of the raw data to the first node. The first node generates a plurality of synthetic data based on the raw data from the second node and the characteristic information of the raw data. In this example, the model training node or the model use node (third node) determines the method for obtaining the characteristic information of the raw data (or the method for obtaining the data label of the raw data) according to the model requirements, which can improve the quality of the synthetic data and thus help improve the performance of model training or updating.

[0331] 8 , which is a flow chart of another model data acquisition method provided by an embodiment of the present application. The example shown in FIG8 is an example of a first node (data synthesis node, or it can be called a central node, a first entity, etc.) and a second node (data providing node, or it can be called a node to be served, a second entity, etc.) as the execution subject of the interaction diagram. The model data acquisition method shown in FIG8 may include steps 801-805. Steps 801-805 are as follows:

[0332] 801. A first node sends seventh information to a second node, where the seventh information indicates a reporting configuration of sixth information. Correspondingly, the second node receives the seventh information.

[0333] For example, the seventh information indicates that the sixth information includes original data; or the seventh information indicates that the sixth information includes original data and data labels of original data; or the seventh information indicates that the sixth information includes original data, data labels of original data, original data priority, etc.

[0334] Optionally, when the seventh information indicates that the sixth information includes the data label of the original data, the first node also sends the eighth information to the second node, and the eighth information indicates the processing method of the second node to obtain the sixth information, and the processing method of the second node to obtain the sixth information includes the method of obtaining the data label of the original data.

[0335] Of course, the method for obtaining the data label of the original data can also be determined by the second node according to the model processing requirements, and this solution does not limit this.

[0336] Furthermore, the second node obtains the data label of the original data based on the acquisition method of the data label of the original data and the original data.

[0337] The second node also reports the data label of the original data to the first node.

[0338] For the introduction of this part, please refer to the description of step 601 in the embodiment shown in FIG6 , which will not be repeated here.

[0339] 802. The second node sends sixth information to the first node, where the sixth information includes the following type of information: original data. Correspondingly, the first node receives the sixth information.

[0340] 803. The second node further sends ninth information to the first node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data. Correspondingly, the first node receives the ninth information.

[0341] That is, in this example, the method for obtaining the characteristic information of the original data is determined by the second node. The second node can determine the method for obtaining the characteristic information of the original data according to the model requirements.

[0342] 804. The first node obtains characteristic information of the original data based on the sixth information and the ninth information.

[0343] The first node can obtain the characteristic information of the original data based on the received original data and the method for obtaining the characteristic information of the original data.

[0344] 805. The first node generates a plurality of synthetic data based on the feature information of the original data and the sixth information. The plurality of synthetic data are used for processing the model.

[0345] For the introduction of this step, please refer to the description of step 602 in the embodiment shown in FIG6 , which will not be repeated here.

[0346] Optionally, the first node is a third-party device that performs the aforementioned actions related to the first node. For example, the above steps 801-805 are all performed by the third-party device.

[0347] Optionally, the second node is a third-party device that performs the aforementioned second-node-related actions. For example, steps 801-803 are performed by the third-party device.

[0348] Optionally, the first node is a network device. For example, the above steps 801-805 are all performed by the network device.

[0349] Optionally, the second node is a terminal device. For example, the above steps 801-803 are performed by the terminal device.

[0350] Optionally, the first node includes a network device and a third-party device. In one example, steps 804 and 805 can be performed by a third-party device, such as an OTT or cloud server, and steps 801-803 can be performed by the network device. Furthermore, the network device and the third-party device can communicate with each other to transmit the content transmitted in steps 801-803.

[0351] In this embodiment of the present application, a first node processes raw data from a second node based on the method for obtaining feature information of the raw data and the method for obtaining feature information of the raw data to obtain feature information of the raw data. The first node then generates multiple synthetic data based on the feature information of the raw data and the raw data. The second node determines the method for obtaining feature information of the raw data based on model requirements, thereby improving the quality of the synthetic data and thereby contributing to improved performance of model training or updating.

[0352] For another example, the following is an introduction using the method in which the third node indicates to the first node the characteristic information of the original data. Referring to Figure 9, it is a flow chart of another model data acquisition method provided by an embodiment of the present application. The example shown in Figure 9 is illustrated by taking the first node (data synthesis node, or it can be called a central node, a first entity, etc.), the second node (data providing node, or it can be called a node to be served, a second entity, etc.) and the third node (model training node or model using node, or it can be called a third entity, such as a model training entity or a model using entity) as an example of the execution subject of the interaction. The model data acquisition method shown in Figure 9 may include steps 901-905. Steps 901-905 are specifically as follows:

[0353] 901. A first node sends seventh information to a second node, where the seventh information indicates a reporting configuration of sixth information. Correspondingly, the second node receives the seventh information.

[0354] For example, the seventh information indicates that the sixth information includes original data; or the seventh information indicates that the sixth information includes original data and data labels of original data; or the seventh information indicates that the sixth information includes original data, data labels of original data, original data priority, etc.

[0355] Optionally, when the seventh information indicates that the sixth information includes the data label of the original data, the first node also sends the eighth information to the second node, and the eighth information indicates the processing method of the second node to obtain the sixth information, and the processing method of the second node to obtain the sixth information includes the method of obtaining the data label of the original data.

[0356] Of course, the method for obtaining the data label of the original data can also be determined by the second node according to the model processing requirements, and this solution does not limit this.

[0357] Furthermore, the second node obtains the data label of the original data based on the acquisition method of the data label of the original data and the original data.

[0358] For the introduction of this part, please refer to the description of step 601 in the embodiment shown in FIG6 , which will not be repeated here.

[0359] 902. The second node sends sixth information to the first node, where the sixth information includes the following type of information: original data. Correspondingly, the first node receives the sixth information.

[0360] 903. The third node sends ninth information to the first node, where the ninth information indicates a method for the first node to obtain the characteristic information of the original data. Correspondingly, the first node receives the ninth information.

[0361] The third node may be a model training node or a model usage node. The model training node or the model usage node determines a feature extraction method for the raw data based on the model requirements, and then sends the feature extraction method for the raw data to the first node.

[0362] In one possible implementation, when the seventh information indicates that the sixth information includes the data tag of the original data, the third node further sends the second node a method for obtaining the data tag of the original data. In other words, the method for obtaining the data tag of the original data may be indicated by the third node to the second node.

[0363] In a possible implementation, the third node further sends the first node a method for obtaining the data label of the original data. That is, the method for obtaining the data label of the original data may be instructed by the third node to the first node.

[0364] 904. The first node obtains feature information of the original data based on the ninth information and the sixth information.

[0365] 905. The first node generates a plurality of synthetic data based on the feature information of the original data and the sixth information. The plurality of synthetic data are used for processing the model.

[0366] For the introduction of this part, please refer to the description of step 602 in the embodiment shown in FIG6 , which will not be repeated here.

[0367] Optionally, the first node is a third-party device that performs the aforementioned actions related to the first node. For example, the above steps 901-905 are all performed by the third-party device.

[0368] Optionally, the second node is a third-party device that performs the aforementioned second-node-related actions. For example, steps 901 and 902 are performed by a third-party device.

[0369] Optionally, the third node is a third-party device that performs the aforementioned third-node-related actions. For example, the above step 903 is performed by the third-party device.

[0370] Optionally, the first node is a network device. For example, the above steps 901-905 are all performed by the network device.

[0371] Optionally, the second node is a terminal device. For example, steps 901 and 902 are performed by the terminal device.

[0372] Optionally, the third node is an AI node. For example, the above step 903 is performed by the AI ​​node.

[0373] Optionally, the first node includes a network device and a third-party device. In one example, steps 904 and 905 can be performed by a third-party device, such as an OTT or cloud server, and steps 901-903 can be performed by the network device. Furthermore, the network device and the third-party device can communicate with each other to transmit the content transmitted in steps 901 and 903.

[0374] Optionally, the third node is an AI entity. In one example, the above step 903 can be performed by a third-party device, such as an OTT or a cloud server.

[0375] In an embodiment of the present application, a first node obtains characteristic information of the original data based on the method for obtaining characteristic information of the original data from the second node and the original data from the third node. The first node then generates multiple synthetic data based on the characteristic information of the original data and the original data. In this example, synthetic data that meets diversity requirements and has the same data distribution as the measured data can be generated based on the original data and the characteristic information of the original data, thereby improving the quality of the synthetic data and thereby helping to improve the performance of model training or updating.

[0376] The following is an introduction using an example in which a third node processes the original data to obtain characteristic information of the original data. Referring to Figure 10, it is a flow chart of another model data acquisition method provided by an embodiment of the present application. Optionally, the method can be applied to the aforementioned model data acquisition system, such as the model data acquisition system shown in Figure 1. The example shown in Figure 10 is illustrated by taking the first node (data synthesis node, or it can be called a central node, a first entity, etc.), the second node (data providing node, or it can be called a node to be served, a second entity, etc.) and the third node (model training node or model using node, or it can be called a third entity, such as a model training entity or a model using entity) as an example of the execution subject of the interaction diagram. The model data acquisition method shown in Figure 10 may include steps 1001-1007. Steps 1001-1007 are specifically as follows:

[0377] 1001. A first node sends seventh information to a second node, where the seventh information indicates a reporting configuration of sixth information. Correspondingly, the second node receives the seventh information.

[0378] For example, the seventh information indicates that the sixth information includes original data; or the seventh information indicates that the sixth information includes original data and data labels of original data; or the seventh information indicates that the sixth information includes original data, data labels of original data, original data priority, etc.

[0379] For the introduction of this part, please refer to the description of step 601 in the embodiment shown in FIG6 , which will not be repeated here.

[0380] 1002. The second node sends sixth information to the first node, where the sixth information includes the following type of information: original data. Correspondingly, the first node receives the sixth information.

[0381] 1003. The second node sends original data to the third node. Correspondingly, the third node receives the original data.

[0382] The second node not only sends the original data to the first node, but also sends the original data to the third node.

[0383] The third node can be a model training node or a model usage node.

[0384] 1004. The third node processes the original data to obtain feature information of the original data.

[0385] The third node (model training node or model use node) determines a method for obtaining the feature information of the original data according to the model requirements, and then processes the received original data based on the method for obtaining the feature information of the original data to obtain the feature information of the original data.

[0386] 1005. The first node sends tenth information to the third node, where the tenth information indicates a reporting configuration for the eleventh information. Correspondingly, the third node receives the tenth information.

[0387] The eleventh information includes the following types of information: characteristic information of the original data. The reporting configuration may include, for example, the characteristic information of the original data and the data volume of the characteristic information of the original data.

[0388] In a possible implementation, the tenth information further indicates a reporting configuration of a data tag of the original data.

[0389] That is to say, the data label of the original data can be reported by the third node or the second node, and this solution does not impose any restrictions on this.

[0390] 1006. The third node sends eleventh information to the first node, where the eleventh information includes the following type of information: characteristic information of the original data. Correspondingly, the first node receives the eleventh information.

[0391] That is, the third node reports the characteristic information of the original data to the first node.

[0392] 1007. The first node generates a plurality of synthetic data based on the sixth information and the eleventh information, and the plurality of synthetic data are used for processing the model.

[0393] For the introduction of the above steps, please refer to the description of step 602 in the embodiment shown in FIG6 , which will not be repeated here.

[0394] Exemplarily, when the eleventh information includes the data label of the original data, in a possible implementation, the first node also sends the eighth information to the second node, and the eighth information indicates the processing method of the second node to obtain the sixth information, and the processing method of the second node to obtain the sixth information includes the method of obtaining the data label of the original data.

[0395] Alternatively, the third node sends eighth information to the second node, where the eighth information indicates a processing method for the second node to obtain the sixth information, and the processing method for the second node to obtain the sixth information includes a method for obtaining a data tag of the original data. This solution does not impose any restrictions on this.

[0396] This example uses the data label of the original data reported by the second node as an example.

[0397] Alternatively, the data label of the raw data may be reported by a third node. Exemplarily, when the eleventh information includes the data label of the raw data, the third node determines a method for obtaining the data label of the raw data based on model requirements. The third node processes the raw data based on the method for obtaining the data label of the raw data to obtain the data label of the raw data. Furthermore, the third node transmits the data label of the raw data to the first node.

[0398] Optionally, the first node is a third-party device that performs the aforementioned actions related to the first node. For example, the above steps 1001, 1002, and 1005-1007 are all performed by a third-party device.

[0399] Optionally, the second node is a third-party device that performs the aforementioned second-node-related actions. For example, the above steps 1001-1003 are performed by the third-party device.

[0400] Optionally, the third node is a third-party device that performs the aforementioned third-node-related actions. For example, the above steps 1003-1006 are performed by the third-party device.

[0401] Optionally, the first node is a network device. For example, the above steps 1001, 1002, and 1005-1007 are all performed by the network device.

[0402] Optionally, the second node is a terminal device. For example, the above steps 1001-1003 are performed by the terminal device.

[0403] Optionally, the third node is an AI node. For example, the above steps 1003-1006 are performed by the AI ​​node.

[0404] Optionally, the first node includes a network device and a third-party device. In one example, step 1007 can be performed by a third-party device, such as an OTT or cloud server, and steps 1001, 1002, 1005, and 1006 can be performed by the network device. Furthermore, the network device and the third-party device can communicate with each other to transmit the content transmitted in steps 1001, 1002, 1005, and 1006.

[0405] Optionally, the third node includes a third-party device, as well as a network device or a terminal device. In one example, step 1004 can be performed by a third-party device, such as an OTT or cloud server, and steps 1003, 1005, and 1006 can be performed by the network device or terminal device. Furthermore, the network device or terminal device can also communicate with the third-party device to transmit the content transmitted in steps 1003, 1005, and 1006.

[0406] In this embodiment of the present application, a first node generates multiple synthetic data based on the raw data from the second node and the feature information of the raw data from the third node. This example generates synthetic data that meets diversity requirements and has the same data distribution as the measured data based on the raw data and the feature information of the raw data, thereby improving the quality of the synthetic data and thereby helping to improve the performance of model training or updating.

[0407] It should be noted that in the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between the various embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.

[0408] It should be noted that the second node and the third node in the embodiment of the present application can be located in the same node or in two different nodes, and this solution does not impose any restrictions on this.

[0409] The above describes in detail the method of the embodiment of the present application, and the following provides the device of the embodiment of the present application. It will be understood that in the various device embodiments of the present application, the division of multiple units or modules is only a logical division based on function, and is not intended to limit the specific structure of the device. In a specific implementation, some functional modules may be subdivided into more small functional modules, and some functional modules may be combined into one functional module, but no matter whether these functional modules are subdivided or combined, the general process performed by the device is the same. For example, some devices include a receiving unit and a sending unit. In some designs, the sending unit and the receiving unit can also be integrated into a communication unit, which can implement the functions implemented by the receiving unit and the sending unit. Typically, each unit corresponds to its own program code (or program instructions), and when the program code corresponding to each of these units runs on the processor, the unit is controlled by the processing unit to execute the corresponding process to implement the corresponding function.

[0410] The embodiments of the present application also provide an apparatus for implementing any of the above methods. For example, a model data acquisition apparatus is provided that includes a module (or means) for implementing each step performed by the first node in any of the above methods. For another example, a model data acquisition apparatus is provided that includes a module (or means) for implementing each step performed by the second node in any of the above methods. For another example, a model data acquisition apparatus is provided that includes a module (or means) for implementing each step performed by the third node in any of the above methods.

[0411] For example, referring to FIG11 , which is a schematic diagram of a structure of a model data acquisition device provided in an embodiment of the present application, the device may include a transceiver module 1101 and a processing module 1102, wherein:

[0412] When the model data acquisition device is used to implement the function of the first node in the above method embodiment, the transceiver module 1101 is used to perform the operation of the first node in step 601 of the embodiment as shown in Figure 6, and the processing module 1102 is used to perform step 602 of the embodiment as shown in Figure 6; or, the transceiver module 1101 is used to perform one or more of the operations of the first node in steps 701 and 704 of the embodiment as shown in Figure 7, and the processing module 1102 is used to perform step 705 of the embodiment as shown in Figure 7; or, the transceiver module 1101 is used to perform one or more of the operations of the first node in steps 801, 802, and 803 of the embodiment as shown in Figure 8, and the processing module 1102 is used to execute one or more of steps 804 and 805 of the embodiment shown in Figure 8; or, the transceiver module 1101 is used to execute one or more of the operations of the first node in steps 901, 902, and 903 of the embodiment shown in Figure 9, and the processing module 1102 is used to execute one or more of the steps 904 and 905 of the embodiment shown in Figure 9; or, the transceiver module 1101 is used to execute one or more of the operations of the first node in steps 1001, 1002, 1005, and 1006 of the embodiment shown in Figure 10, and the processing module 1102 is used to execute step 1007 of the embodiment shown in Figure 10.

[0413] When the model data acquisition device is used to implement the function of the second node in the above-mentioned method embodiment, the transceiver module 1101 is used to perform the operation of the second node in step 601 of the embodiment as shown in Figure 6; or, the transceiver module 1101 is used to perform one or more of the operations of the second node in steps 701, 702, and 704 of the embodiment as shown in Figure 7, and the processing module 1102 is used to perform step 703 of the embodiment as shown in Figure 7; or, the transceiver module 1101 is used to perform one or more of the operations of the second node in steps 801, 802, and 803 of the embodiment as shown in Figure 8; or, the transceiver module 1101 is used to perform one or more of the operations of the second node in steps 901 and 902 of the embodiment as shown in Figure 9; or, the transceiver module 1101 is used to perform one or more of the operations of the second node in steps 1001, 1002, and 1003 of the embodiment as shown in Figure 10.

[0414] When the model data acquisition device is used to implement the function of the third node in the above-mentioned method embodiment, the transceiver module 1101 is used to perform the operation of the third node in step 702 of the embodiment as shown in Figure 7; or, the transceiver module 1101 is used to perform the operation of the third node in step 903 of the embodiment as shown in Figure 9; or, the transceiver module 1101 is used to perform one or more of the operations of the third node in steps 1003, 1005, and 1006 of the embodiment as shown in Figure 10, and the processing module 1102 is used to perform step 1004 of the embodiment as shown in Figure 10.

[0415] The detailed description of the operations involved in each of the above modules can be found in the description of the above embodiments, which will not be repeated here.

[0416] It should be understood that the division of the modules in the above devices is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. In addition, the modules in the model data acquisition device can be implemented in the form of a processor calling software; for example, the model data acquisition device includes a processing circuit, the processing circuit is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of the modules of the device, wherein the processing circuit is a processor or a part of the processing circuit in the processor, the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory inside the device or a memory outside the device. Alternatively, the modules in the device can be implemented in the form of hardware circuits, and the functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units by designing the logical relationship of the components in the circuit. For another example, in another implementation, the hardware circuit can be implemented by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units. All modules of the above devices can be implemented in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.

[0417] FIG12 is a schematic diagram of the hardware structure of another model data acquisition device provided in an embodiment of the present application. The model data acquisition device 1200 shown in FIG12 includes one or more processing circuits 1201 (one processing circuit is shown in the figure). The processing circuit 1201 can be one or more processors, or a circuit used for processing in one or more processors.

[0418] Optionally, the model data acquisition device 1200 may further include a transceiver circuit 1202 (indicated by a dotted line in the figure). The processing circuit 1201 and the transceiver circuit 1202 are coupled to each other. The transceiver circuit 1202 may be a transceiver or an interface circuit. For example, when the device 1200 is a network device, a terminal device, a core network device, or an AI entity, the transceiver circuit 1202 may be a transceiver or an interface circuit; when the device 1200 is a chip for a network device, a terminal device, a core network device, or an AI entity, the transceiver circuit 1202 may be an interface circuit. For example, the AI ​​entity may be a third-party device, such as an OTT, or a cloud server.

[0419] Optionally, the model data acquisition device 1200 may further include a memory 1203 (indicated by a dotted line in the figure). The memory 1203 is used to store instructions executed by the processing circuit 1201, or to store input data required by the processing circuit 1201 to execute instructions, or to store data generated after the processing circuit 1201 executes instructions.

[0420] Optionally, the memory 1203 may be located in the one or more processors, or located outside the one or more processors, or may include a storage part located in the one or more processors and a storage part located outside the one or more processors.

[0421] The memory 1203 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM).

[0422] The memory 1203 can store programs. When the program stored in the memory 1203 is executed by the processing circuit 1201, the processing circuit 1201 and the transceiver circuit 1202 are used to execute the various steps of the model data acquisition method of the embodiment of the present application.

[0423] Processing circuit 1201 is a circuit capable of processing signals. In one implementation, processing circuit 1201 may be a circuit capable of reading and executing instructions, such as one or more of the following processors: a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP), or a processing circuit within the aforementioned processors. In another implementation, processing circuit 1201 may implement certain functions through the logical relationships of hardware circuits, where the logical relationships of the hardware circuits are fixed or reconfigurable. For example, processing circuit 1201 is one or more of the following processors: a hardware circuit implemented by an ASIC or a programmable logic device (PLD), such as an FPGA, or a processing circuit within the aforementioned processors. In a reconfigurable hardware circuit, the process of the processor loading a configuration file to implement hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC or a processing circuit in an ASIC, such as one or more of the following processors: a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc., or a processing circuit in the aforementioned processor. The processing circuit 1201 is used to execute relevant programs to implement the functions required to be performed by the units in the model data acquisition device of the embodiment of the present application, or to execute the model data acquisition method of the method embodiment of the present application.

[0424] It can be seen that each module in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms or part of the processing circuits in these processors.

[0425] In addition, the modules in the above device can be fully or partially integrated together, or can be implemented independently. In one implementation, these modules are integrated together and implemented in the form of a system-on-a-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the modules of the device. The type of the at least one processor can be different, for example, including a CPU and FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0426] The transceiver circuit 1202 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the device 1200 and other devices or a communication network. For example, information can be obtained through the transceiver circuit 1202.

[0427] It should be noted that although the device 1200 shown in FIG12 only shows a processing circuit, a transceiver circuit, and a memory, during the specific implementation process, those skilled in the art will understand that the device 1200 also includes other components necessary for normal operation. At the same time, according to specific needs, those skilled in the art will understand that the device 1200 may also include hardware components that implement other additional functions. In addition, those skilled in the art will understand that the device 1200 may also include only the components necessary to implement the embodiments of the present application, and does not necessarily include all the components shown in FIG12.

[0428] An embodiment of the present application further provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is executed on a computer or a processor, the computer or processor executes one or more steps in any of the above methods.

[0429] The present application also provides a computer program product comprising instructions, which, when executed on a computer or processor, causes the computer or processor to execute one or more steps in any of the above methods.

[0430] It should be understood that in the description of this application, unless otherwise specified, " / " indicates that the objects associated with each other are in an "or" relationship. For example, A / B can mean A or B; where A and B can be singular or plural. Also, in the description of this application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural. In addition, to facilitate the clear description of the technical solutions of the embodiments of this application, in the embodiments of this application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily mean different. At the same time, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0431] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling, direct coupling, or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, and can be electrical, mechanical or other forms.

[0432] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0433] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic medium such as a floppy disk, a hard disk, a tape, a magnetic disk, or an optical medium such as a digital versatile disc (DVD), or a semiconductor medium such as a solid state disk (SSD).

[0434] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for obtaining model data, characterized in that, Executed by a first node or a circuit for the first node, the method includes: Receiving first information from a second node, the first information including information of the following types: original data and feature information of the original data; based on the first information, generating a plurality of synthetic data for processing by a model.

2. The method according to claim 1, wherein The method further includes: Sending second information to the second node, the second information indicating a processing manner of the second node for obtaining the first information, and the processing manner of the second node for obtaining the first information including a manner of obtaining the feature information of the original data.

3. The method according to claim 2, characterized in that, The first information further includes information of the following types: a data label of the original data and / or a priority of the original data.

4. The method according to claim 3, characterized in that, The processing manner of the second node for obtaining the first information further includes a manner of obtaining the data label of the original data.

5. The method according to any one of claims 1 to 4, characterized in that The method further includes: Sending third information to the second node, the third information indicating a reporting configuration of the first information.

6. The method according to claim 5, characterized in that, The reporting configuration includes at least one of the following: a type of the first information, a data volume of the first information.

7. The method according to any one of claims 1 to 6, characterized in that The method further includes: Sending fourth information to the second node, the fourth information indicating that the first node has an ability to jointly process feature information of multiple original data in the same group.

8. The method according to any one of claims 1 to 7, characterized in that, The first information further includes information of the following types: group information of the original data.

9. The method according to any one of claims 1 to 8, characterized in that The method further includes: Sending at least one subset of the plurality of synthetic data to the second node; Receiving fifth information from the second node, the fifth information indicating an evaluation result of at least one synthetic data in the at least one subset, and the evaluation result indicating at least one synthetic data to be excluded in the at least one subset, or indicating at least one synthetic data not to be excluded in the at least one subset.

10. The method according to claim 9, characterized in that, The evaluation result is used to screen the plurality of synthetic data to obtain at least one screened synthetic data; wherein, the at least one screened synthetic data is used for processing by the model.

11. The method according to claim 10, characterized in that, The at least one screened synthetic data is determined based on a distance between a first synthetic data to be excluded and other synthetic data in the plurality of synthetic data, and the first synthetic data to be excluded is a synthetic data to be excluded determined based on the fifth information.

12. A method for obtaining model data, characterized in that, Executed by a first node or a circuit for the first node, the method includes: Receiving sixth information from a second node, the sixth information including information of the following types: original data; Based on the sixth information and the feature information of the original data, generating a plurality of synthetic data for processing by a model.

13. The method according to claim 12, wherein The method further includes: Receiving eleventh information from a third node, the eleventh information including information of the following types: the feature information of the original data.

14. The method according to claim 12, wherein The method further includes: Receiving ninth information from the second node, the ninth information indicating a manner of obtaining the feature information of the original data by the first node; Processing the original data based on the ninth information to obtain the feature information of the original data.

15. The method according to claim 12, characterized in that, The method further includes: Receive the ninth information from the third node, where the ninth information indicates the acquisition method for the first node to obtain the feature information of the original data; process the original data based on the ninth information to obtain the feature information of the original data.

16. The method according to claim 12, characterized in that, The feature information of the original data is determined based on the acquisition method for the feature information of the original data, and the acquisition method for the feature information of the original data is determined by the first node.

17. The method according to any one of claims 12 to 16, characterized in that, The sixth information further includes at least one of the following types of information: the priority of the original data, or the data label of the original data.

18. The method according to any one of claims 12 to 17, characterized in that The method further includes: Send the eighth information to the second node, where the eighth information indicates the processing method for the second node to obtain the sixth information, and the processing method for the second node to obtain the sixth information includes the acquisition method for the data label of the original data.

19. The method according to any one of claims 12 to 18, characterized in that, The method further includes: Send the seventh information to the second node, where the seventh information indicates the reporting configuration for the sixth information.

20. The method according to any one of claims 12 to 19, characterized in that, The reporting configuration includes at least one of the following: the type of the sixth information, or the data volume of the sixth information.

21. The method according to any one of claims 12 to 20, characterized in that, The method further includes: Send the fourth information to the second node, where the fourth information indicates that the first node has the ability to jointly process the feature information of multiple original data in the same group.

22. The method according to any one of claims 12 to 21, characterized in that, The sixth information further includes the following type of information: the group information of the original data.

23. The method according to any one of claims 12 to 22, characterized in that, The method further includes: Send at least one subset of the multiple synthesized data to the second node; Receive the fifth information from the second node, where the fifth information indicates the evaluation result of at least one synthesized data in the at least one subset, and the evaluation result indicates at least one synthesized data that needs to be excluded in the at least one subset, or indicates at least one synthesized data that does not need to be excluded in the at least one subset.

24. The method according to claim 23, wherein The evaluation result is used to screen the multiple synthesized data to obtain at least one screened synthesized data; wherein, the at least one screened synthesized data is used for the processing of the model.

25. The method according to claim 24, characterized in that, The at least one screened synthesized data is determined based on the distance between the first synthesized data to be excluded and other synthesized data in the multiple synthesized data, and the first synthesized data to be excluded is the synthesized data that needs to be excluded determined based on the fifth information.

26. A method for obtaining model data, characterized in that Executed by the second node or a circuit for the second node, the method includes: Send the first information to the first node, where the first information includes the following types of information: the original data and the feature information of the original data.

27. The method according to claim 26, wherein The method further includes: Receive the second information from the first node, where the second information indicates the processing method for the second node to obtain the first information, and the processing method includes the acquisition method for the feature information of the original data; Obtain the first information based on the second information.

28. The method according to claim 27, wherein The first information further includes the following types of information: the data label of the original data and / or the priority of the original data.

29. The method according to any one of claims 26 to 28, characterized in that, The method further includes: Receive the third information from the first node, where the third information indicates the reporting configuration for the first information.

30. The method according to claim 29, characterized in that, The reported configuration includes at least one of the following: the type of the first information, the data volume of the first information.

31. The method according to any one of claims 26 to 30, characterized in that, The method further includes: Receiving fourth information from the first node, the fourth information indicating that the first node has the ability to jointly process the characteristic information of multiple pieces of original data in the same group.

32. The method according to any one of claims 26 to 31, characterized in that, The first information further includes information of the following type: group information of the original data.

33. The method according to any one of claims 26 to 32, characterized in that, The method further includes: Receiving at least one subset of multiple synthesized data from the first node; Sending fifth information to the first node, the fifth information indicating the evaluation result of at least one synthesized data in the at least one subset; the evaluation result indicates at least one synthesized data to be excluded in the at least one subset, or indicates at least one synthesized data that does not need to be excluded in the at least one subset.

34. The method according to claim 33, wherein The evaluation result is determined based on the distance between at least one synthesized data in the at least one subset and local data.

35. The method according to claim 33 or 34, characterized in that, The evaluation result is used to screen the multiple synthesized data to obtain at least one screened synthesized data; wherein, the at least one screened synthesized data is used for model processing.

36. A method for obtaining model data, characterized in that, Executed by a second node or a circuit for the second node, the method includes: Sending sixth information to the first node, the sixth information including information of the following type: original data; Sending ninth information to the first node, the ninth information indicating the acquisition method for the first node to obtain the characteristic information of the original data.

37. The method according to claim 36, wherein The method further includes: Receiving eighth information from the first node, the eighth information indicating the processing method for the second node to obtain the sixth information, and the processing method for the second node to obtain the sixth information includes the acquisition method of the data label of the original data.

38. The method according to claim 36 or 37, characterized in that, The method further includes: Receiving seventh information from the first node, the seventh information indicating the reporting configuration of the sixth information.

39. The method according to any one of claims 36 to 38, characterized in that, The method further includes: Receiving fourth information from the first node, the fourth information indicating that the first node has the ability to jointly process the characteristic information of multiple pieces of original data in the same group.

40. The method according to any one of claims 36 to 39, characterized in that, The method further includes: Receiving at least one subset of multiple synthesized data from the first node; Sending fifth information to the first node, the fifth information indicating the evaluation result of at least one synthesized data in the at least one subset, the evaluation result indicating at least one synthesized data to be excluded in the at least one subset, or indicating at least one synthesized data that does not need to be excluded in the at least one subset.

41. The method according to claim 40, characterized in that, The evaluation result is determined based on the distance between at least one synthesized data in the at least one subset and local data.

42. The method according to claim 40 or 41, characterized in that, The evaluation result is used to screen the multiple synthesized data to obtain at least one screened synthesized data; wherein, the at least one screened synthesized data is used for model processing.

43. A method for obtaining model data, characterized in that, Executed by a third node or a circuit for the third node, the method includes: Sending eleventh information to the first node, the eleventh information including information of the following type: characteristic information of the original data.

44. A method for obtaining model data, characterized in that, Executed by a third node or a circuit for the third node, the method includes: Send the ninth information to the first node, where the ninth information indicates the acquisition method for the first node to obtain the feature information of the original data.

45. A model data acquisition device, characterized in that, It includes modules for implementing the method according to any one of claims 1-44.

46. A communication device, characterized in that, The device includes a processor and a memory. The memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1-44.

47. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-44 is implemented.

48. A computer program product, characterized in that, The computer program product includes the program instructions involved. When the involved program instructions are executed, the method according to any one of claims 1-44 is implemented.

Citation Information

Patent Citations

  • Model data acquisition method, device and system

    CN120234608A

  • Feature engineering arrangement method and device

    CN113159145A

  • Data model training method and device

    CN114548416A

  • Model training method and related device

    CN115423031A

  • Communication method and device

    CN116318481A