Data processing method, neural network model training method and related apparatus

WO2026200491A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/082050
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-06
Publication Date
2026-10-01

Smart Images

  • Figure CN2026082050_01102026_PF_FP_ABST
    Figure CN2026082050_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a data processing method, a neural network model training method and a related apparatus. The data processing method comprises: acquiring first historical data to be inferred and first data to be inferred, and on the basis of the first historical data to be inferred and the first data to be inferred, obtaining a first data sequence; on the basis of a label of the first historical data to be inferred and an initial label of the first data to be inferred, obtaining a first label sequence, wherein labels in the first label sequence correspond to data in the first data sequence on a one-to-one basis in terms of sequence positions; and on the basis of the first data sequence, obtaining first association weights between data in the first historical data to be inferred and data in the first data to be inferred, and obtaining an inference result of the first data to be inferred in combination with the first association weights and the first label sequence. The embodiments of the present application facilitate a reduction in the training or fine-tuning overheads of an AI model, and reduce the training latency of the AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing methods, neural network model training methods, and related devices

[0001] This application claims priority to Chinese Patent Application No. 202510381030.X, filed on March 26, 2025, entitled "Data Processing Method, Training Method for Neural Network Model and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of communication technology, and in particular to a data processing method, a neural network model training method, and related apparatus. Background Technology

[0003] Wireless communication systems consist of numerous cells, as shown in Figure 1. The distribution of scatterers and transmission settings often differ across different cellular scenarios, resulting in a highly heterogeneous electromagnetic environment. In the field of machine learning (AI), data from specific scenarios is typically used to train the corresponding AI model for inference within that scenario. Even in transfer learning, data from new scenarios is used to fine-tune the trained AI model. However, with so many cellular scenarios, training or fine-tuning the AI ​​model individually for each scenario increases hardware overhead and training latency. Summary of the Invention

[0004] This application provides a data processing method, a neural network model training method, and related apparatus, which helps to solve the problem that AI models need to be trained or fine-tuned across different scenarios, saves the training or fine-tuning costs of AI models, and reduces the training latency of AI models.

[0005] Firstly, embodiments of this application provide a data processing method applied to an electronic device. Unless otherwise specified, the electronic device in this application can refer to the electronic device itself (e.g., a server, network device, terminal device, Operation Administration and Maintenance (OAM) device, etc.) or a module within the electronic device. For example, the module can be a processing module, a communication module, or a circuit or chip responsible for processing or communication functions within the electronic device, such as a modem chip (also known as a baseband chip), or a system-on-a-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip. Alternatively, it can be a logic module or software capable of implementing all or part of the functions of the electronic device. For ease of description, the following uses an electronic device as an example. The method includes:

[0006] The electronic device acquires first historical data to be inferred and first data to be inferred, and obtains a first data sequence based on the first historical data to be inferred and the first data to be inferred. Based on the labels of the first historical data to be inferred and the initial labels of the first data to be inferred, the electronic device can obtain a first label sequence, wherein the labels in the first label sequence correspond one-to-one with the data in the first data sequence in terms of sequence position. Based on the first data sequence, the electronic device can obtain the first association weight between each data in the first historical data to be inferred and the first data to be inferred, and combine the first association weight and the first label sequence to obtain the inference result corresponding to the first data to be inferred.

[0007] As can be seen, in this embodiment, the electronic device can obtain the first historical data to be inferred and the first association weight between each data in the first data sequence based on the first data sequence, and obtain the inference result corresponding to the first data to be inferred by combining the first association weight and the first label sequence corresponding to the first data sequence. Since the first association weight can characterize the causal relationship, correlation or change pattern between each data in the first data to be inferred and the first historical data to be inferred, the electronic device can infer the first data to be inferred based on this association relationship between the data and the labels corresponding to the historical data, thereby obtaining the inference result corresponding to the first data to be inferred. This eliminates the need to use data from the scene to which the first data to be inferred belongs to specifically train or fine-tune the AI ​​model for inference, which helps to avoid the various hardware and software overheads brought about by training or fine-tuning the AI ​​model. The inference of the first data to be inferred does not have to wait for the retraining or fine-tuning of the AI ​​model, saving the intermediate training latency. In addition, even if the electronic device obtains the reasoning result corresponding to the first data to be reasoned through the trained neural network model, the neural network model can also obtain the reasoning result corresponding to the first data to be reasoned based on the first association weight and the first label sequence, without the need to fine-tune the neural network model. This can also achieve the effect of reducing overhead and saving latency, and can also realize the cross-scene generalization of the neural network model.

[0008] In conjunction with the first aspect, in one possible implementation, the electronic device can perform neighborhood sampling on the first data to be inferred in the data feature space to obtain the first historical data to be inferred.

[0009] In this implementation, a certain amount of first historical data to be inferred can be obtained by sampling the neighborhood of the first data to be inferred. The first historical data to be inferred in the neighborhood usually has a high similarity with the first data to be inferred in the feature space, which can reflect the local structure, trend and pattern of the data. Using these data and the first data to be inferred to construct an input data sequence is helpful for electronic devices to understand the meaning of the first data to be inferred in a specific context or scenario, thereby improving the inference accuracy based on the first data to be inferred.

[0010] In conjunction with the first aspect, in one possible implementation, the electronic device can also acquire the estimated reasoning result of the first data to be reasoned, and perform neighborhood sampling on the estimated reasoning result in the label feature space to obtain the first target label. The historical data to be reasoned corresponding to the first target label is the first historical data to be reasoned.

[0011] In this implementation, the electronic device can acquire the estimated inference result of the first data to be inferred, and perform neighborhood sampling on the estimated inference result in the label feature space to obtain a first target label with a high similarity to the estimated inference result. Similar labels usually also have first historical data to be inferred that have a certain degree of similarity to the first data to be inferred. Using these data and the first data to be inferred to construct an input data sequence helps the electronic device understand the meaning of the first data to be inferred in a specific context or scenario, thereby improving the inference accuracy based on the first data to be inferred.

[0012] In conjunction with the first aspect, in one possible implementation, the electronic device can obtain a first feature tensor of the first label sequence based on the first label sequence, and then combine the first association weight and the first feature tensor to obtain the reasoning result corresponding to the first data to be reasoned.

[0013] In this implementation, the first feature tensor can reflect the high-dimensional and complex interactions and temporal patterns between tags in the tag sequence. Combining this feature tensor with the first association weight can provide the electronic device with richer and more comprehensive information, so as to help the electronic device understand the association between the data in the first tag sequence more accurately, thereby improving the reasoning accuracy based on the first data to be reasoned.

[0014] In conjunction with the first aspect, in one possible implementation, the electronic device can obtain the reasoning result corresponding to the first data to be reasoned based on the first association weights and the first label sequence using a trained neural network model. The neural network model includes N encoder layers, each encoder layer comprising a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The input to the second module in the first encoder layer is a tensor obtained based on the first label sequence. The electronic device processes this tensor through the second module in the first encoder layer to obtain the first value V1 tensor. The output tensor of the second module in the (i-1)th encoder layer serves as the input to the second module in the ith encoder layer. The electronic device processes the output tensor of the second module in the (i-1)th encoder layer through the second module in the ith encoder layer to obtain the ith V1 tensor, where 1 < i ≤ N. When i equals N, the electronic device processes the output tensor of the second module in the (N-1)th encoder layer through the second module in the Nth encoder layer to obtain the Nth V1 tensor. The first feature tensor of the first label sequence includes these N V1 tensors.

[0015] In this implementation, the electronic device processes the first label sequence through a linear layer to obtain the input tensor of the second module in the first encoder layer. This second module then processes the input tensor to obtain the first V1 tensor. Based on this first V1 tensor and the first attention weight, the electronic device obtains the output tensor of the second module in the first encoder layer. For subsequent encoder layers, the second module, based on the output tensor of the second module in the previous encoder layer, obtains the V1 tensor extracted by the second module in the current encoder layer. Based on this, the neural network model with N encoder layers will extract N V1 tensors. These N V1 tensors contain relevant information from the first label sequence. Combining this information with the corresponding attention weights for attention calculation provides additional guidance to the neural network model, helping it focus on labels related to the first data to be inferred, thereby improving the model's inference performance on that data.

[0016] In conjunction with the first aspect, in one possible implementation, the electronic device can obtain a tensor corresponding to the first data sequence based on the first data sequence. This tensor is then processed by a first module in the first encoder layer to obtain a first query Q tensor and a first key K tensor. Based on this first Q tensor and the first K tensor, the electronic device can obtain a first attention weight. This first attention weight and the first V1 tensor are then processed by a second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The electronic device then processes the output tensor of the first module in the (i-1)th encoder layer using the first module in the i-th encoder layer to obtain an i-th Q tensor and an i-th K tensor. Based on the i-th Q tensor and the i-th K tensor, the electronic device can obtain an i-th attention weight. This i-th attention weight and the i-th V1 tensor are then processed by a second module in the i-th encoder layer to obtain the output tensor of the second module in the i-th encoder layer. When i equals N, the electronic device can obtain N attention weights, and the first association weight includes these N attention weights. Based on the output tensor of the second module in the Nth encoder layer, the electronic device can obtain the final inference result corresponding to the first data to be inferred.

[0017] In this implementation, based on the connection between the first and second modules in each encoder layer, the electronic device can obtain the attention weights generated in each encoder layer. It then processes these attention weights and the V1 tensor generated by the second module in the current encoder layer to obtain the output tensor of the second module in the current encoder layer. When the output tensor of the second module in the Nth encoder layer is obtained, the electronic device processes this output tensor through a linear layer to obtain the inference result corresponding to the first data to be inferred. For each encoder layer, the second module obtains the attention weights through the connection with the first module and combines the corresponding V1 tensor with the attention weights. This introduces label information to guide the neural network model inference to obtain the output tensor of the second module in the current encoder layer, which helps improve the neural network model's understanding of the first data to be inferred and enhances the model's interpretability.

[0018] In conjunction with the first aspect, in one possible implementation, the first module in the first encoder layer processes the tensor obtained based on the first data sequence to obtain the first V2 tensor. The electronic device then processes the first Q tensor, the first K tensor, and the first V2 tensor through the first module in the first encoder layer to obtain the output tensor of the first module in the first encoder layer. Similarly, the first module in the i-th encoder layer processes the output tensor of the first module in the (i-1)-th encoder layer to obtain the i-th V2 tensor. The electronic device then processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor through the first module in the i-th encoder layer to obtain the output tensor of the first module in the i-th encoder layer.

[0019] In this implementation, the first module in each encoder layer focuses on the attention mechanism processing at the data level. The Q tensor, K tensor, and V2 tensor generated by the first module in the current encoder layer are used to generate the output tensor of the first module, so as to provide input for the first module in the next encoder layer, thereby enabling the next encoder layer to generate attention weights based on the input.

[0020] In conjunction with the first aspect, in one possible implementation, the first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer. The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the second module in the i-th encoder layer.

[0021] In this implementation, attention weights in each encoder layer can be obtained either by the first module based on the Q-tensor and K-tensor generated by the first module, or by the second module based on the Q-tensor and K-tensor generated by the first module. The generation of attention weights is relatively flexible due to the different connection methods between the first and second modules.

[0022] In conjunction with the first aspect, in one possible implementation, the electronic device can obtain the reasoning result corresponding to the first data to be reasoned based on the first association weights and the first label sequence using a trained neural network model. The neural network model includes an encoder layer, which comprises a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The first feature tensor includes a V1 tensor. Obtaining the first feature tensor based on the first label sequence involves processing the tensor obtained based on the first label sequence through the second module in the encoder layer to obtain the V1 tensor.

[0023] In this implementation, the electronic device processes the first label sequence through a linear layer to obtain the input tensor of the second module in the encoder layer. The second module processes the input tensor to obtain the V1 tensor, which contains relevant information in the first label sequence. Combining this information with attention weights to perform attention calculation can provide additional guidance for the neural network model.

[0024] In conjunction with the first aspect, in one possible implementation, the first association weight includes attention weights. The reasoning result of the first data to be reasoned is obtained based on the first association weights and the first feature tensor, including: processing the attention weights and the V1 tensor through a second module in the encoder layer to obtain the output tensor of the second module in the encoder layer; and obtaining the reasoning result corresponding to the first data to be reasoned based on the output tensor of the second module in the encoder layer. The attention weights are obtained based on the Q tensor and the K tensor, which are obtained by processing the tensor based on the first data sequence by the first module in the encoder layer.

[0025] In this implementation, based on the connection between the first module and the second module in the encoder layer, the electronic device can obtain the attention weights generated in the encoder layer, and process the attention weights and the V1 tensor generated by the second module in the encoder layer to obtain the output tensor of the second module in the encoder layer. Based on the output tensor, the reasoning result of the first data to be reasoned can be obtained.

[0026] In conjunction with the first aspect, in one possible implementation, the attention weights are obtained by the first module in the encoder layer processing the Q tensor and the K tensor, or the attention weights are obtained by the second module in the encoder layer processing the Q tensor and the K tensor.

[0027] In this implementation, the attention weights in the encoder layer can be obtained by either the first module based on the Q-tensor and K-tensor generated by the first module, or by the second module based on the Q-tensor and K-tensor generated by the first module. The generation of attention weights is relatively flexible due to the different connection methods between the first and second modules.

[0028] Secondly, embodiments of this application provide a method for training a neural network model, applied to an electronic device. Unless otherwise specified, the electronic device in this application can refer to the electronic device itself (e.g., a server, network device, terminal device, OAM device, etc.) or a module within the electronic device. For example, the module can be a processing module, a communication module, or a circuit or chip responsible for communication functions within the electronic device, such as a modem chip (also known as a baseband chip), or a SoC chip or SIP chip containing a modem core. Alternatively, it can be a logic module or software capable of implementing all or part of the functions of the electronic device. It should be understood that the electronic device here can be the electronic device described in the first aspect above, or it may not be the electronic device described in the first aspect above. For ease of description, the following uses an electronic device as an example, and the method includes:

[0029] The electronic device acquires second historical data to be inferred and second data to be inferred, and obtains a second data sequence based on the second historical data to be inferred and the second data to be inferred. Based on the labels of the second historical data to be inferred and the initial labels of the second data to be inferred, the electronic device can obtain a second label sequence, wherein the labels in the second label sequence correspond one-to-one with the data in the second data sequence in terms of sequence position. The electronic device performs the following processing through a neural network: based on the second data sequence, it obtains a second association weight between each data in the second historical data to be inferred and the second data to be inferred; combining the second association weight and the second label sequence, it can obtain the inference result corresponding to the second data to be inferred. Based on the inference result and the labels of the second data to be inferred, the electronic device can adjust the parameters of the neural network, and obtain a trained neural network model after multiple rounds of iterative training.

[0030] As can be seen, in this embodiment, the electronic device can obtain the second correlation weight between the second historical data to be inferred and the data in the second data sequence based on the second data sequence, and combine the second correlation weight with the second label sequence corresponding to the second data sequence to obtain the inference result corresponding to the second data to be inferred. Since the second correlation weight can characterize the causal relationship, correlation or change pattern between the data in the second data to be inferred and the data in the second historical data to be inferred, the electronic device can infer the second data to be inferred based on this correlation relationship between the data and the labels corresponding to the historical data, thereby obtaining the inference result corresponding to the second data to be inferred. Based on the inference result of the second data to be inferred and its label, the electronic device can adjust the parameters of the neural network to obtain a neural network model that can be directly applied to various scenarios for data inference. Since the neural network learns to infer the current data based on the correlation relationship between the data and the labels of the historical data during the training phase, the neural network model does not need to be specially trained or fine-tuned using the data of that scenario when applied to different scenarios. This helps to avoid the various hardware and software overheads brought about by training or fine-tuning the neural network model, and can be directly used for inference of scenario data, which helps to save the intermediate training or fine-tuning latency and realize the cross-scenario generalization of the neural network model.

[0031] In conjunction with the second aspect, in one possible implementation, the electronic device can perform neighborhood sampling on the second data to be inferred in the data feature space to obtain the second historical data to be inferred.

[0032] In this implementation, a certain number of second historical data to be inferred can be obtained by sampling the neighborhood of the second data to be inferred. The second historical data to be inferred in the neighborhood usually has a high similarity with the second data to be inferred in the feature space, which can reflect the local structure, trend and pattern of the data. Using these data and the second data to be inferred to construct an input data sequence is helpful for electronic devices to understand the meaning of the second data to be inferred in a specific context or scenario, thereby improving the inference accuracy based on the second data to be inferred.

[0033] In conjunction with the second aspect, in one possible implementation, the electronic device can also perform neighborhood sampling on the label of the second data to be inferred in the label feature space to obtain the second target label, and the historical data to be inferred corresponding to the second target label is the second historical data to be inferred.

[0034] In this implementation, the electronic device can perform neighborhood sampling on the labels of the second data to be inferred in the label feature space to obtain a second target label with high similarity to the first label. Similar labels usually also have second historical data to be inferred that have a certain degree of similarity to the second data to be inferred. Using these data and the second data to be inferred to construct an input data sequence helps the electronic device understand the meaning of the second data to be inferred in a specific context or scenario, thereby improving the inference accuracy based on the second data to be inferred.

[0035] In conjunction with the second aspect, in one possible implementation, the electronic device can obtain a second feature tensor of the second label sequence based on the second label sequence, and then combine the second association weight and the second feature tensor to obtain the reasoning result corresponding to the second data to be reasoned.

[0036] In this implementation, the second feature tensor can reflect the high-dimensional and complex interactions and temporal patterns between labels in the label sequence. Combining this feature tensor with the second association weight can provide electronic devices with richer and more comprehensive information, which can help electronic devices understand the association between data in the second data sequence more accurately, thereby improving the reasoning accuracy based on the second data to be reasoned.

[0037] In conjunction with the second aspect, in one possible implementation, the neural network includes N encoder layers. Each encoder layer includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The input to the second module in the first encoder layer is a tensor obtained based on the second label sequence. The electronic device processes this tensor through the second module in the first encoder layer to obtain the first value V1 tensor. The output tensor of the second module in the (i-1)th encoder layer is used as the input to the second module in the ith encoder layer. The electronic device processes the output tensor of the second module in the (i-1)th encoder layer through the second module in the ith encoder layer to obtain the ith V1 tensor, where 1 < i ≤ N. When i equals N, the electronic device processes the output tensor of the second module in the (N-1)th encoder layer through the second module in the Nth encoder layer to obtain the Nth V1 tensor. The first feature tensor of the first label sequence includes these N V1 tensors.

[0038] In this implementation, the electronic device processes the second label sequence through a linear layer to obtain the input tensor of the second module in the first encoder layer. This second module then processes the input tensor to obtain the first V1 tensor. Based on this first V1 tensor and the first attention weight, the electronic device obtains the output tensor of the second module in the first encoder layer. For subsequent encoder layers, the second module, based on the output tensor of the second module in the previous encoder layer, obtains the V1 tensor extracted by the second module in the current encoder layer. Thus, a neural network with N encoder layers extracts N V1 tensors. These N V1 tensors contain relevant information from the second label sequence. Combining this information with the corresponding attention weights for attention calculation provides additional guidance to the neural network, helping it focus on labels related to the second data to be inferred, thereby improving the neural network's inference performance on the second data.

[0039] In conjunction with the second aspect, in one possible implementation, the electronic device can obtain a tensor corresponding to the second data sequence based on the second data sequence. This tensor is then processed by the first module in the first encoder layer to obtain the first query Q tensor and the first key K tensor. Based on the first Q tensor and the first K tensor, the electronic device can obtain the first attention weight. This first attention weight and the first V1 tensor are then processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The electronic device then processes the output tensor of the first module in the (i-1)th encoder layer using the first module in the i-th encoder layer to obtain the i-th Q tensor and the i-th K tensor. Based on the i-th Q tensor and the i-th K tensor, the electronic device can obtain the i-th attention weight. This i-th attention weight and the i-th V1 tensor are then processed by the second module in the i-th encoder layer to obtain the output tensor of the second module in the i-th encoder layer. When i equals N, the electronic device can obtain N attention weights, and the second association weight includes these N attention weights. Based on the output tensor of the second module in the Nth encoder layer, the electronic device can obtain the final inference result of the second data to be inferred.

[0040] In this implementation, based on the connection between the first and second modules in each encoder layer, the electronic device can obtain the attention weights generated in each encoder layer. It then processes these attention weights and the V1 tensor generated by the second module in the current encoder layer to obtain the output tensor of the second module in the current encoder layer. Upon obtaining the output tensor of the second module in the Nth encoder layer, the electronic device processes this output tensor through a third Linear layer to obtain the inference result for the second data to be inferred. For each encoder layer, the second module obtains the attention weights through the connection with the first module and combines the corresponding V1 tensor with the attention weights. This introduces label information to guide the neural network inference to obtain the output tensor of the second module in the current encoder layer, which helps improve the neural network's understanding of the second data to be inferred and enhances the interpretability of the neural network.

[0041] In conjunction with the second aspect, in one possible implementation, the first module in the first encoder layer processes the tensor obtained based on the second data sequence to obtain the first V2 tensor. The electronic device then processes the first Q tensor, the first K tensor, and the first V2 tensor through the first module in the first encoder layer to obtain the output tensor of the first module in the first encoder layer. Similarly, the first module in the i-th encoder layer processes the output tensor of the first module in the (i-1)-th encoder layer to obtain the i-th V2 tensor. The electronic device then processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor through the first module in the i-th encoder layer to obtain the output tensor of the first module in the i-th encoder layer.

[0042] In this implementation, the first module in each encoder layer focuses on the attention mechanism processing at the data level. The Q tensor, K tensor, and V2 tensor generated by the first module in the current encoder layer are used to generate the output tensor of the first module, so as to provide input for the first module in the next encoder layer, thereby enabling the next encoder layer to generate attention weights based on the input.

[0043] In conjunction with the second aspect, in one possible implementation, the first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer. The ith attention weight is obtained by processing the ith Q tensor and the ith K tensor through the first module in the ith encoder layer, or the ith attention weight is obtained by processing the ith Q tensor and the ith K tensor through the second module in the ith encoder layer.

[0044] In this implementation, attention weights in each encoder layer can be obtained either by the first module based on the Q-tensor and K-tensor generated by the first module, or by the second module based on the Q-tensor and K-tensor generated by the first module. The generation of attention weights is relatively flexible due to the different connection methods between the first and second modules.

[0045] In conjunction with the second aspect, in one possible implementation, the neural network includes an encoder layer, which comprises a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The second feature tensor includes a V1 tensor. Obtaining the second feature tensor based on the second label sequence involves processing the tensor obtained based on the second label sequence through the second module in the encoder layer to obtain the V1 tensor.

[0046] In this implementation, the electronic device processes the second label sequence through a linear layer to obtain the input tensor of the second module in the encoder layer. The second module processes the input tensor to obtain the V1 tensor, which contains relevant information in the second label sequence. Combining this information with attention weights to perform attention calculation can provide additional guidance for the neural network.

[0047] In conjunction with the second aspect, in one possible implementation, the second association weight includes attention weights. The reasoning result for the second data to be reasoned is obtained based on the second association weights and the second feature tensor, including: processing the attention weights and the V1 tensor through a second module in the encoder layer to obtain the output tensor of the second module in the encoder layer; and obtaining the reasoning result corresponding to the second data to be reasoned based on the output tensor of the second module in the encoder layer. The attention weights are obtained based on the Q tensor and the K tensor, which are obtained by processing the tensor based on the second data sequence by the first module in the encoder layer.

[0048] In this implementation, based on the connection between the first module and the second module in the encoder layer, the electronic device can obtain the attention weights generated in the encoder layer, and process the attention weights and the V1 tensor generated by the second module in the encoder layer to obtain the output tensor of the second module in the encoder layer. Based on the output tensor, the reasoning result of the second data to be reasoned can be obtained.

[0049] In conjunction with the second aspect, in one possible implementation, the attention weights are obtained by the first module in the encoder layer processing the Q tensor and the K tensor, or the attention weights are obtained by the second module in the encoder layer processing the Q tensor and the K tensor.

[0050] In this implementation, the attention weights in the encoder layer can be obtained by either the first module based on the Q-tensor and K-tensor generated by the first module, or by the second module based on the Q-tensor and K-tensor generated by the first module. The generation of attention weights is relatively flexible due to the different connection methods between the first and second modules.

[0051] Thirdly, embodiments of this application provide a data processing apparatus, which includes modules or units for performing the method described in the first aspect, such as a first acquisition unit and a first processing unit; wherein:

[0052] The first acquisition unit is used to acquire a first data sequence and a first tag sequence. The first data sequence includes first historical data to be inferred and first data to be inferred. The tags in the first tag sequence correspond one-to-one with the data in the first data sequence in terms of sequence position.

[0053] The first processing unit is used to obtain the first association weight between each data in the first data to be inferred and the first historical data to be inferred based on the first data sequence; and to obtain the inference result corresponding to the first data to be inferred based on the first association weight and the first label sequence.

[0054] In conjunction with the third aspect, in one possible implementation, the first processing unit is further configured to: perform neighborhood sampling on the first data to be inferred in the data feature space to obtain the first historical data to be inferred.

[0055] In conjunction with the third aspect, in one possible implementation, the first acquisition unit is further configured to: acquire the estimated reasoning result of the first data to be reasoned; the first processing unit is further configured to: perform neighborhood sampling on the estimated reasoning result in the label feature space to obtain the first target label, and the first historical data to be reasoned is the historical data to be reasoned corresponding to the first target label.

[0056] In conjunction with the third aspect, in one possible implementation, regarding the reasoning result of obtaining the first data to be reasoned based on the first association weight and the first label sequence, the first processing unit is specifically used for:

[0057] The first feature tensor of the first label sequence is obtained based on the first label sequence;

[0058] The reasoning result corresponding to the first data to be reasoned is obtained based on the first association weight and the first feature tensor.

[0059] In conjunction with the third aspect, in one possible implementation, the reasoning result of the first data to be reasoned based on the first association weight and the first label sequence is obtained through a neural network model. The neural network model includes N encoder layers, and each of the N encoder layers includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism.

[0060] The first feature tensor includes N value V1 tensors. In obtaining the first feature tensor of the first label sequence based on the first label sequence, the first processing unit is specifically used for:

[0061] The tensor obtained based on the first label sequence is processed by the second module in the first encoder layer of N encoder layers to obtain the first V1 tensor among N V1 tensors;

[0062] The output tensor of the second module in the (i-1)th encoder layer of N encoder layers is processed by the second module in the i-th encoder layer to obtain the i-th V1 tensor among N V1 tensors, where 1 < i ≤ N.

[0063] In conjunction with the third aspect, in one possible implementation, the first association weight includes N attention weights. Regarding obtaining the reasoning result of the first data to be reasoned based on the first association weight and the first feature tensor, the first processing unit is specifically used for:

[0064] The first attention weight and the first V1 tensor out of N attention weights are processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The first attention weight is obtained based on the first query Q tensor and the first key K tensor. The first Q tensor and the first K tensor are obtained by the first module in the first encoder layer processing the tensor obtained based on the first data sequence.

[0065] The second module in the i-th encoder layer processes the i-th attention weight and the i-th V1 tensor among the N attention weights to obtain the output tensor of the second module in the i-th encoder layer. The i-th attention weight is obtained based on the i-th Q tensor and the i-th K tensor. The i-th Q tensor and the i-th K tensor are obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0066] When i equals N, the inference result corresponding to the first data to be inferred is obtained based on the output tensor of the second module in the Nth encoder layer.

[0067] In conjunction with the third aspect, in one possible implementation, the first processing unit is further configured to:

[0068] The first module in the first encoder layer processes the first Q tensor, the first K tensor, and the first V2 tensor to obtain the output tensor of the first module in the first encoder layer. The first V2 tensor is obtained by the first module in the first encoder layer processing the tensor based on the first data sequence.

[0069] The first module in the i-th encoder layer processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor to obtain the output tensor of the first module in the i-th encoder layer. The i-th V2 tensor is obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0070] In conjunction with the third aspect, in one possible implementation, the first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer.

[0071] The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the second module in the i-th encoder layer.

[0072] It should be understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the first aspect of the embodiments of this application should be adapted to the third aspect of the embodiments of this application simultaneously, and can achieve the same or similar beneficial effects, which will not be repeated here.

[0073] Fourthly, embodiments of this application provide a training apparatus for a neural network model, the apparatus including modules or units for performing the method described in the second aspect above, such as a second acquisition unit and a second processing unit; wherein:

[0074] The second acquisition unit is used to acquire a second data sequence and a second tag sequence. The second data sequence includes second historical data to be reasoned and second data to be reasoned. The tags in the second tag sequence correspond one-to-one with the data in the second data sequence in terms of sequence position.

[0075] The second processing unit is used to perform the following processing through a neural network: obtaining the second association weight between each data in the second data to be inferred and the second historical data to be inferred based on the second data sequence, and obtaining the inference result corresponding to the second data to be inferred based on the second association weight and the second label sequence;

[0076] The second processing unit is also used to adjust the parameters of the neural network based on the reasoning result of the second data to be reasoned and the label of the second data to be reasoned, so as to obtain a neural network model.

[0077] In conjunction with the fourth aspect, in one possible implementation, the second processing unit is further configured to: perform neighborhood sampling on the second data to be inferred in the data feature space to obtain the second historical data to be inferred.

[0078] In conjunction with the fourth aspect, in one possible implementation, the second processing unit is further configured to: perform neighborhood sampling on the labels of the second data to be inferred in the label feature space to obtain the second target label, and the second historical data to be inferred is the historical data to be inferred corresponding to the second target label.

[0079] In conjunction with the fourth aspect, in one possible implementation, regarding obtaining the reasoning result corresponding to the second data to be reasoned based on the second association weight and the second label sequence, the second processing unit is specifically used for:

[0080] The second feature tensor of the second label sequence is obtained based on the second label sequence;

[0081] The reasoning result corresponding to the second data to be reasoned is obtained based on the second association weight and the second feature tensor.

[0082] In conjunction with the fourth aspect, in one possible implementation, the neural network includes N encoder layers, each of the N encoder layers including a first module based on a self-attention mechanism and a second module based on a self-attention mechanism;

[0083] The second feature tensor comprises N V1 tensors. Regarding the acquisition of the second feature tensor from the second label sequence based on the second label sequence, the second processing unit is specifically used for:

[0084] The tensor obtained based on the second label sequence is processed by the second module in the first encoder layer of N encoder layers to obtain the first V1 tensor among N V1 tensors;

[0085] The output tensor of the second module in the (i-1)th encoder layer of N encoder layers is processed by the second module in the i-th encoder layer to obtain the i-th V1 tensor among N V1 tensors, where 1 < i ≤ N.

[0086] In conjunction with the fourth aspect, in one possible implementation, the second association weight includes N attention weights. Regarding obtaining the inference result corresponding to the second data to be inferred based on the second association weight and the second feature tensor, the second processing unit is specifically used for:

[0087] The first attention weight and the first V1 tensor out of N attention weights are processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The first attention weight is obtained based on the first query Q tensor and the first key K tensor. The first Q tensor and the first K tensor are obtained by the first module in the first encoder layer processing the tensor obtained based on the second data sequence.

[0088] The second module in the i-th encoder layer processes the i-th attention weight and the i-th V1 tensor among the N attention weights to obtain the output tensor of the second module in the i-th encoder layer. The i-th attention weight is obtained based on the i-th Q tensor and the i-th K tensor. The i-th Q tensor and the i-th K tensor are obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0089] When i equals N, the reasoning result corresponding to the second data to be reasoned is obtained based on the output tensor of the second module in the Nth encoder layer.

[0090] In conjunction with the fourth aspect, in one possible implementation, the second processing unit is further configured to:

[0091] The first module in the first encoder layer processes the first Q tensor, the first K tensor, and the first V2 tensor to obtain the output tensor of the first module in the first encoder layer. The first V2 tensor is obtained by the first module in the first encoder layer processing the tensor based on the second data sequence.

[0092] The first module in the i-th encoder layer processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor to obtain the output tensor of the first module in the i-th encoder layer. The i-th V2 tensor is obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0093] In conjunction with the fourth aspect, in one possible implementation, the first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer.

[0094] The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the second module in the i-th encoder layer.

[0095] It should be understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the second aspect of the embodiments of this application should be adapted to the fourth aspect of the embodiments of this application simultaneously, and can achieve the same or similar beneficial effects, which will not be repeated here.

[0096] Fifthly, embodiments of this application provide an electronic device for implementing any one of the first or second aspects described above, or for implementing the method in any implementation of any one of the first or second aspects described above. For example, the device may be a data processing device, a module applied to a data processing device (e.g., a processor, chip, or chip system), or a logic node, logic module, or software capable of implementing all or part of the functions of a data processing device. For example, the device may be a training device for a neural network model, a module applied to a training device for a neural network model (e.g., a processor, chip, or chip system), or a logic node, logic module, or software capable of implementing all or part of the functions of a training device for a neural network model.

[0097] In one possible implementation, the electronic device in the fifth aspect above includes units, modules, or means for performing the methods in any one or any implementation of the first or second aspect above. Specifically, the units, modules, or means may be implemented in software, in hardware, or in a combination of software and hardware.

[0098] In another possible implementation, the electronic device in the fifth aspect above includes at least one processor; the at least one processor is configured to perform the corresponding functions in the data processing method or the training method of the neural network model described above.

[0099] Optionally, the at least one processor may be coupled to at least one memory for storing programs (instructions) and / or data (such as one or more computer programs) necessary for the device. Optionally, the electronic device may further include a communication interface for enabling communication between the electronic device and other devices. Optionally, the at least one memory may be located internally or externally to the electronic device.

[0100] Optionally, the electronic device may further include a transceiver, with the processor coupled to the transceiver. The processor executes computer programs or instructions to control the transceiver to receive and transmit information. When the processor executes the computer programs or instructions, it is also used to implement the above method through logic circuits or executed code instructions. The transceiver can be a transceiver circuit, a transceiver module, or an input / output interface, used to receive signals from other devices outside the electronic device and transmit them to the processor, or to send signals from the processor to other devices outside the electronic device. When the electronic device is a chip, the transceiver is a transceiver circuit or an input / output interface.

[0101] When the electronic device in the fifth aspect above is a chip, the transmitting unit can be an output unit, such as an output circuit or a communication interface; the receiving unit can be an input unit, such as an input circuit or a communication interface. When the electronic device is a physical device, the transmitting unit can be a transmitter or a receiver; the receiving unit can be a receiver or a receiver.

[0102] In a sixth aspect, embodiments of this application provide a chip, including: a processor, configured to call and run a computer program from a memory, causing a device / apparatus on which the chip is mounted to perform the method as described in any of the embodiments of the first or second aspect above.

[0103] In a seventh aspect, embodiments of this application provide a computer-readable storage medium storing a computer program for execution by a device / apparatus, wherein the computer program, when executed, implements the method as described in any of the embodiments of the first or second aspect above.

[0104] Eighthly, embodiments of this application provide a computer program product that, when run by a device, causes the device to perform the method as described in any of the embodiments of the first or second aspect above. Attached Figure Description

[0105] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.

[0106] Figure 1 is a schematic diagram of a cellular scenario provided in an embodiment of this application;

[0107] Figure 2 is a schematic diagram of a communication system provided in an embodiment of this application;

[0108] Figure 2a is a schematic diagram of another communication system provided in an embodiment of this application;

[0109] Figure 3 is a schematic diagram of another communication system provided in an embodiment of this application;

[0110] Figure 3a is a schematic diagram of another communication system provided in an embodiment of this application;

[0111] Figure 4 is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0112] Figure 5 is a schematic diagram of neighborhood sampling in a cellular scenario provided by an embodiment of this application;

[0113] Figure 6 is a flowchart illustrating another data processing method provided in an embodiment of this application;

[0114] Figure 7 is a schematic diagram of the structure of a neural network model provided in an embodiment of this application;

[0115] Figure 8 is a schematic diagram of an encoder layer provided in an embodiment of this application;

[0116] Figure 9 is a schematic diagram of another encoder layer provided in an embodiment of this application;

[0117] Figure 10 is a flowchart illustrating a training method for a neural network model provided in an embodiment of this application;

[0118] Figure 11 is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0119] Figure 12 is a schematic diagram of the structure of a training device for a neural network model provided in an embodiment of this application;

[0120] Figure 13 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0121] Figure 14 is a schematic diagram of a baseband hardware provided in an embodiment of this application. Detailed Implementation

[0122] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0123] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0124] The terms “component,” “module,” “system,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, an application running on a terminal device and the terminal device can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0125] AI can endow machines with human-like intelligence, for example, allowing them to use computer hardware and software to simulate certain intelligent human behaviors. To achieve artificial intelligence, machine learning methods can be employed. In machine learning, machines learn (or train) a model using training data. This model represents the mapping between inputs and outputs. The learned model can be used for reasoning (or prediction), that is, it can be used to predict the output corresponding to a given input. This output can also be called the reasoning result (or prediction result).

[0126] This document explains some basic concepts in the field of AI, which does not limit the scope of protection of the embodiments of this application.

[0127] (1) Machine learning (ML)

[0128] Machine learning is an important technological approach to achieving AI. AI encompasses machine learning and other methods. Machine learning refers to learning models or rules from raw data, such as neural networks, decision trees, and support vector machines. Machine learning can be categorized into supervised learning, unsupervised learning, and reinforcement learning.

[0129] Supervised learning, based on collected sample values ​​and labels, uses machine learning algorithms to learn the mapping relationship between sample values ​​and labels, and expresses this learned mapping relationship using a machine learning model. The process of training the machine learning model is the process of learning this mapping relationship. For example, in signal detection, the noisy received signal is the sample, and the corresponding real constellation point is the label. Machine learning aims to learn the mapping relationship between samples and labels through training, that is, to enable the machine learning model to learn a signal detector. During training, the model parameters are optimized by calculating the error between the model's predicted values ​​and the real labels. Once the mapping relationship is learned, it can be used to predict the sample label of each new sample. The mapping relationship learned in supervised learning can include linear mappings and nonlinear mappings. Based on the type of label, the learning task can be divided into classification tasks and regression tasks.

[0130] Unsupervised learning relies solely on collected sample values, using algorithms to discover inherent patterns within the samples. One type of unsupervised learning algorithm uses the samples themselves as supervisory signals; that is, the model learns the mapping relationship from sample to sample, which is called self-supervised learning. During training, model parameters are optimized by calculating the error between the model's predictions and the samples themselves. Self-supervised learning can be used for signal compression and decompression recovery applications; common algorithms include autoencoders and generative adversarial networks.

[0131] Reinforcement learning, unlike supervised learning, is a type of algorithm that learns problem-solving strategies through interaction with the environment. Unlike supervised and unsupervised learning, reinforcement learning problems do not have explicit "correct" action labels. The algorithm needs to interact with the environment to obtain reward signals from the environment, and then adjust its decision actions to obtain a larger reward signal value. For example, in downlink power control, the reinforcement learning model adjusts the downlink transmission power of each terminal device based on the total system throughput feedback from the wireless network, aiming to achieve a higher system throughput. The goal of reinforcement learning is also to learn the mapping relationship between the environment state and the optimal decision action. However, because the label of the "correct action" cannot be obtained in advance, the network cannot be optimized by calculating the error between the action and the "correct action." Reinforcement learning training is achieved through iterative interaction with the environment.

[0132] Deep neural networks (DNNs) are a specific implementation of machine learning. According to the general approximation theorem, neural networks can theoretically approximate any continuous function, thus enabling them to learn arbitrary mappings. Traditional communication systems rely on extensive expert knowledge to design communication modules, while DNN-based deep learning communication systems can automatically discover hidden pattern structures from large datasets, establish mapping relationships between data, and achieve performance superior to traditional modeling methods.

[0133] Based on their construction method, DNNs can be divided into feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs). FNNs can be neural networks where neurons in adjacent layers are completely connected pairwise, which makes FNNs typically require a large amount of storage space and have high computational complexity.

[0134] CNNs are neural networks specifically designed to process data with a grid-like structure. For example, time-series data (discrete sampling along the time axis) and image data (two-dimensional discrete sampling) can both be considered grid-like data. CNNs do not use all the input information at once for computation; instead, they use a fixed-size window to extract a portion of the information for convolution operations, which significantly reduces the computational cost of model parameters. Furthermore, depending on the type of information extracted by the window (such as people and objects in an image representing different types of information), each window can use different convolution kernels, allowing CNNs to better extract features from the input data.

[0135] Recurrent Neural Networks (RNNs) are a type of distributed neural network (DNN) that utilizes feedback time-series information. Their input includes the current input value and their own output value from the previous time step. RNNs are well-suited for acquiring temporally correlated sequence features, and are particularly applicable to applications such as speech recognition and channel coding / decoding.

[0136] AI models refer to function models that map inputs of a certain dimension to outputs of a certain dimension, and their parameters can be obtained through machine learning training. For example, f(x) = ax 2 +b is a quadratic function model, which can be viewed as an AI model. a and b correspond to the parameters of this model and can be obtained through machine learning training. In machine learning, the data used for model training, validation, and / or testing can form a dataset or training dataset. The quantity and / or quality of data in the dataset or training dataset will affect the effectiveness of machine learning. Model training involves selecting an appropriate loss function (which measures the difference between the model's predictions and the true values) and using optimization algorithms to train the model parameters to minimize the loss function value. Model testing involves evaluating the model's performance using test data after training. Model application involves using the trained model to solve real-world problems.

[0137] A neural network, or artificial neural network, is a mathematical model that mimics the behavioral characteristics of animal neural networks to perform distributed parallel information processing; it is a special form of AI model. AI models can also be referred to as AI functions.

[0138] (2) Model Training

[0139] Model training involves selecting an appropriate function (such as a loss function) and using optimization algorithms to train the model parameters so that the difference between the model's predicted values ​​and the ground truth (or target values, labels) tends to be minimized.

[0140] For example, model training methods include, but are not limited to, supervised learning, self-supervised learning, and knowledge distillation.

[0141] (3) Model files and model parameters

[0142] Model files and / or model parameters can be used to determine the model. Optionally, the model in this application may refer to the model itself, or it may refer to the model files and / or model parameters used to determine the model.

[0143] The model file can be used to indicate the model structure, which may include, but is not limited to, FNN, CNN, or RNN. The model file can have a fixed format, such as a standard predefined format, or a format pre-negotiated by both ends of the interface. Model parameters can refer to parameters in the neural network model, such as, but not limited to, the number of layers in the neural network, the type and weights of neurons in each layer, etc. This application does not limit the method of distributing model parameters.

[0144] Take DNN as an example. The idea behind DNN comes from the neuronal structure of the brain. Each neuron can perform a weighted summation operation on its inputs and then use the result of the weighted summation operation to generate the output through a non-linear function. For example, the input of a neuron is x = [x0, x1, ..., x...]. N-1 The weights corresponding to the inputs are w = [w0, w1, ..., w] N-1 The bias of the weighted summation is b. The nonlinear function f() can take many forms; for example, the nonlinear function f() can be the maximum value function max{0, x}. Then the effect of a neuron's execution is... Where N is a positive integer, and n is a positive integer greater than or equal to 0 and less than or equal to (N-1). The weights of the weighted summation operation of neurons in a neural network and the nonlinear function are called the parameters of the neural network. The parameters of all neurons in a neural network constitute the parameters of the neural network.

[0145] A DNN typically has multiple neural network layers, including an input layer, one or more hidden layers, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Each layer contains multiple neurons. Layers are fully connected; that is, any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. The input layer processes the received values ​​(i.e., the DNN's input) through neurons and then passes them to the hidden layers. Similarly, the hidden layers pass the computation results to the final output layer, producing the DNN's output. This application does not limit the structure and parameters used in the AI ​​model.

[0146] One of the model structure or model parameters can be predefined, while the other can be provided by the sender (e.g., the network side). Alternatively, both the model structure and model parameters can be provided by the sender (e.g., the network side). This application does not impose any limitations on this.

[0147] The transmitting model can refer to sending model files and / or model parameters, while the receiving model can refer to receiving model files and / or model parameters. Currently, AI technology is being introduced into wireless communication systems. AI technology can be used for wireless channel information compression and reconstruction, beam management, and positioning enhancement, improving wireless communication performance based on trained AI models. A wireless AI framework can include multiple modules such as data collection, model training, model management, model inference, and model storage.

[0148] (4) Generalization of AI models: This refers to the ability of a model to reason correctly or make appropriate judgments when it is reasoning on new data that has not been trained on. Generalization is a key indicator for measuring the performance of a model in practical applications, especially when dealing with unseen data that is different from the training set.

[0149] (5) Transfer learning: In traditional machine learning, it is usually necessary to train a model from scratch and rely on a large amount of labeled data for learning. However, in transfer learning, the model or features learned on one task can be transferred to another task, thereby reducing the need for a large amount of labeled data and improving learning efficiency.

[0150] (6) Meta-learning: also known as "learning how to learn," unlike traditional machine learning methods that rely on large amounts of data to train a model to solve a specific task, meta-learning focuses on how to enable the model to learn from a small amount of data and quickly adapt to new environments or tasks. One of the goals of meta-learning is to enable the model to quickly adjust its internal representation and parameters after seeing a small number of new task samples in order to perform inference as quickly and accurately as possible.

[0151] While transfer learning can reduce the training time of AI models, it requires fine-tuning the parameters of the source model using data from the target task. Limited learning, a key direction in meta-learning, aims to enable models to learn effectively with a very small number of training samples. In other words, meta-learning still requires training data to fine-tune the source model, helping it quickly adapt to new environments or tasks based on past experience. Training or fine-tuning with new data in new scenarios incurs significant overhead in data acquisition, labeling, transmission, computation, and storage, increasing training and usage latency. Therefore, achieving cross-scenario generalization of AI models is a pressing issue that needs to be addressed.

[0152] To overcome the shortcomings of existing technologies, this application provides a data processing method, a neural network model training method, and related apparatus. This data processing method and neural network model training method can be implemented based on the communication system shown in Figure 2. As shown in Figure 2, this communication system includes terminal devices and network devices. Communication scenarios between terminal devices, between terminal devices and network devices, and between network devices generate a large amount of wireless communication data, such as signal propagation data, radio spectrum data, network topology data, environmental characteristic data, signal quality and performance data, user behavior and load data, etc. This data can be used for training AI models or for inference or reasoning of specific tasks. In some possible implementations, devices outside the communication system shown in Figure 2 can perform AI model training or AI task inference based on the data generated in the communication system. In another possible implementation, one or more devices within the communication system shown in Figure 2 can perform AI model training or AI task inference based on the data generated in the communication system. That is, AI model training or AI task inference can be performed on one device or jointly on multiple devices.

[0153] For example, the communication systems shown in Figure 2 include, but are not limited to: Narrow Band-Internet of Things (NB-IoT), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access 2000 (CDMA2000), Time Division-Synchronization Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), 5th generation (5G) mobile communication systems, future evolution systems, or multiple communication convergence systems. For example, the communication system shown in Figure 2 can be applied to machine-to-machine (M2M), enhanced mobile broadband (eMBB), ultra-reliable and low-latency communication (uRLLC), massive machine-type communication (mMTC) and other scenarios.

[0154] Optionally, the communication system may also include at least one AI node.

[0155] Figure 2a is another schematic diagram of a communication system applicable to an embodiment of this application. Compared to the communication system shown in Figure 1, the communication system 100 shown in Figure 2a further includes an AI node 140. The AI ​​node 140 is used to perform AI-related operations, such as building training datasets, training or inferring AI models, etc.

[0156] In one implementation, network device 110 can send data related to AI model training to AI node 140, whereby AI node 140 constructs a training dataset and trains the AI ​​model. As an example, the data related to AI model training may include data reported by terminal device 120 or terminal device 130. AI node 140 can send the results of AI model-related operations to network device 110, and then forward them to terminal device 120 or terminal device 130 via network device 110. For example, the results of AI model-related operations may include at least one of the following: a trained AI model, model evaluation results, or test results, etc. Exemplarily, a portion of the trained AI model may be deployed on network device 110, and another portion on terminal device 120 or terminal device 130. Optionally, the trained AI model may be deployed on network device 110, or it may be deployed on terminal device 120 or terminal device 130.

[0157] It should be understood that Figure 2a is only used as an example of AI node 140 being directly connected to network device 110. In other scenarios, AI node 140 can also be connected to terminal device 120 or terminal device 130. Alternatively, AI node 140 can be connected to network device 110, terminal device 120, and terminal device 130 simultaneously. Alternatively, AI node 140 can also be connected to one or more of network device 110, terminal device 120, and terminal device 130 through a third-party network element. This application embodiment does not limit the connection relationship between AI network element and other network elements.

[0158] Alternatively, in another implementation, the AI ​​node 140 can also be set as a module in network devices and / or terminal devices.

[0159] For example, the data processing method and neural network model training method provided in this application can also be implemented based on the communication system shown in Figure 3. As shown in Figure 3, the communication system includes terminal equipment, access network equipment, core network equipment, and Operation Administration and Maintenance (OAM). The network elements are connected through interfaces (e.g., next generation (NG) interfaces, Xn interfaces, F1 interfaces) or air interfaces. The access network equipment can be a single radio access network (RAN) node, or it can include multiple RAN nodes, such as a central unit (CU) and a distributed unit (DU). The CU can also be divided into CU-control plane (CP) and CU-user plane (UP), or CU can also be called open (O)-CU, DU can also be called O-DU, CU-CP can also be called O-CU-CP, and CU-UP can also be called O-CU-UP. Figure 3 shows that each network element can deploy one or more AI modules (Figure 3 uses one AI module as an example). These AI modules are used to implement corresponding AI functions, and can then be used to implement the data processing method and / or neural network model training method provided in this application. The AI ​​modules deployed in different network elements can be the same or different. Different AI modules can be understood as the AI ​​module models being configured with different parameters so that different AI modules can achieve different functions. The AI ​​module model can be configured based on one or more of the following parameters: structural parameters (e.g., at least one of the following: number of neural network layers, neural network width, inter-layer connections, neuron weights, neuron activation function, and bias in the activation function), input parameters (e.g., type and / or dimension of input parameters), and output parameters (e.g., type and / or dimension of output parameters). The bias in the activation function can also be called the neural network bias. For example, one AI module in Figure 3 can deploy one or more models. The learning process, training process, and inference process of different models can be deployed on different nodes or devices, or they can be deployed on the same node or device.

[0160] The network device can be a network device equipped with one or more AI modules. For example, the network device can be one or more devices in the core network, access network, or OAM as shown in Figure 3. The AI ​​module can be the RAN intelligent controller (RIC) shown in Figure 3a, such as a near real-time RIC or a non-real-time RIC. For example, a near real-time RIC is set in a RAN node (e.g., in a CU or DU), while a non-real-time RIC is set in the OAM, cloud server, core network device, or other network device.

[0161] Figure 3a illustrates another possible application framework in a communication system. As shown in Figure 3a, the communication system includes a Resource Interchange (RIC). For example, an RIC can be an AI module in a network device, used to implement AI-related functions. RICs include near-real-time RICs (near-RT RICs) and non-real-time RICs (non-RT RICs). Non-real-time RICs primarily process non-real-time information, such as data that is not sensitive to latency, with latency in the order of seconds. Real-time RICs primarily process near-real-time information, such as data that is relatively sensitive to latency, with latency in the order of tens of milliseconds.

[0162] Near real-time RICs are used for model training and inference. For example, they are used to train AI models and then use those models for inference. Near real-time RICs can obtain network-side and / or terminal-side information from RAN devices (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminal devices. This information can be used as training data or as data for inference.

[0163] Optionally, near real-time RIC can deliver inference results to RAN devices and / or terminal devices.

[0164] Optionally, inference results can be exchanged between CU and DU, and / or between DU and RU. For example, near real-time RIC submits inference results to DU, and DU sends them to RU.

[0165] Non-real-time RICs are also used for model training and inference. For example, they can be used to train AI models and then use those models for inference. Non-real-time RICs can obtain network-side and / or terminal-side information from RAN devices (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or as inference data, and the inference results can be delivered to RAN nodes and / or terminals.

[0166] Optionally, inference results can be exchanged between CU and DU, and / or between DU and RU. For example, a non-real-time RIC can submit inference results to DU, which in turn can send them to RU.

[0167] Near real-time RICs and non-real-time RICs can also be configured as separate devices. Alternatively, near real-time RICs and non-real-time RICs can also be part of other devices. For example, near real-time RICs can be configured in RAN nodes (e.g., CU, DU), while non-real-time RICs can be configured in OAM, cloud servers, core network devices, or other devices.

[0168] In the embodiments of this application, the terminal device may also be referred to as user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user apparatus. The terminal device can be a device that provides voice / data, such as a handheld device or vehicle-mounted device with wireless connectivity. Currently, examples of terminals include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, wearable devices, terminal devices in 5G networks, or terminal devices in future communication systems, etc., and this application does not limit these examples.

[0169] In this embodiment, the device used to implement the functions of the terminal device can be the terminal device itself, or any device capable of supporting the terminal device in implementing corresponding functions, such as a processor, circuit, or chip. This device can be configured in the terminal device or used in conjunction with the terminal device. In this embodiment, the terminal device is used as an example to illustrate the function of the terminal device, and this does not constitute a limitation on the solution of this embodiment.

[0170] The network devices in this application embodiment may include radio access network (RAN) nodes that connect terminal devices to wireless networks, such as base stations. Base stations can broadly encompass various names as follows, or be replaced by the following names: NodeB, evolved NodeB (eNB), next-generation NodeB (gNB), relay station, access point, transmitting and receiving point (TRP), transmitting point (TP), master station, auxiliary station, motor slide retainer (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), etc. A base station can be a macro base station, micro base station, relay node, donor node, or a combination thereof. A base station can also refer to a communication module, modem, or chip installed within the aforementioned equipment or apparatus. A base station can also be a mobile switching center, equipment performing base station functions in D2D, V2X, and M2M communications, or equipment performing base station functions in future communication systems. A base station can support networks using the same or different access technologies. Optionally, a RAN node can also be a server, wearable device, vehicle, or in-vehicle equipment. For example, the access network equipment in vehicle-to-everything (V2X) technology can be a roadside unit (RSU). The embodiments of this application do not limit the specific technologies or equipment forms used in the network equipment.

[0171] Base stations can be fixed or mobile. For example, a helicopter or drone can be configured to act as a mobile base station, and one or more cells can move depending on the location of the mobile base station. In other examples, a helicopter or drone can be configured as a device to communicate with another base station.

[0172] The technical solution provided in this application will be described in detail below with reference to specific implementation methods.

[0173] Please refer to Figure 4, which is a flowchart illustrating a data processing method provided in an embodiment of this application. This method can be executed by an electronic device. As shown in Figure 4, the method includes steps 401-403:

[0174] 401: Obtain the first data sequence and the first label sequence. The first data sequence includes the first historical data to be inferred and the first data to be inferred. The labels in the first label sequence correspond one-to-one with the data in the first data sequence in terms of sequence position.

[0175] In this embodiment, both the first data sequence and the first label sequence can be referred to as input sequences. If the electronic device performs reasoning through a trained neural network model to obtain the reasoning result of the first data to be reasoned, then the first data sequence and the first label sequence can be referred to as the input sequences of the neural network model. For ease of distinction, the first data sequence can be represented as the input sequence Seq_H, and the first label sequence can be represented as the input sequence Seq_x. Of course, the letters H and x are merely representations; different letters can be used to represent the data sequence and label sequence in different scenarios. The first historical data to be reasoned and the first data to be reasoned can both be referred to as input data, and the labels in the first label sequence can be referred to as input labels. If the electronic device performs reasoning through a neural network model to obtain the reasoning result of the first data to be reasoned, then the first historical data to be reasoned and the first data to be reasoned can be referred to as the input data of the neural network model, and the labels in the first label sequence can be referred to as the input labels of the neural network model. Correspondingly, the reasoning result of the first data to be reasoned can be referred to as the output of the electronic device. If the electronic device performs reasoning through a neural network model to obtain the reasoning result of the first data to be reasoned, then the reasoning result of the first data to be reasoned can be further referred to as the output of the neural network model.

[0176] For example, the first historical data to be inferred and the first inference data can also be wireless channel data, signal propagation data, radio spectrum data, network topology data, environmental feature data, or signal quality data, etc. For instance, when the first historical data to be inferred and the first inference data are environmental feature data, the corresponding tag or inference result can be the path loss of signal transmission. It should be understood that the specific type of data is merely for illustrative purposes and does not impose any limitation on the embodiments of this application.

[0177] It should be noted that the first data to be inferred is the data that needs to be input and inferred at present, while the first historical data to be inferred includes data that has been input and inferred in the past and / or data in the training set. There can be one or more first historical data sets, each with a corresponding label. For example, in a scenario where user location is performed using Channel State Information (CSI), the first data sequence can be represented as follows: in, to Let S represent historical CSIs, n represent the number of historical CSIs, j distinguish different historical CSIs, t represent the position (or index) of the corresponding historical CSI (or historical data) in a set or sequence, and H represent the CSI that needs to be input into the model for inference. Each historical CSI in the sequence Seq_H has a corresponding label, i.e., the corresponding position information. Therefore, the first label sequence can be represented as follows: The position of each CSI in the first data sequence corresponds to the position of its label in the first label sequence. Where x init The initial label represents the inference result corresponding to the first data H to be inferred. This initial label can be a randomly initialized label, or it can be the maximum, minimum, or average value of the labels of the first historical data to be inferred; this application does not impose any limitations on this.

[0178] For example, one implementation of obtaining a first data sequence and a first label sequence is as follows: neighborhood sampling is performed on the first data to be inferred in the data feature space to obtain a preset number of first historical data to be inferred; the preset number of first historical data to be inferred and the first data to be inferred are combined to form a first data sequence; and the labels of the first historical data to be inferred and the initial labels corresponding to the first data to be inferred are combined to form a first label sequence. For example, the data feature space can be Euclidean space, spherical space, manifold space, etc. Neighborhood sampling is performed in the data feature space to select first historical data to be inferred whose feature distance to the first data to be inferred H is less than or equal to a preset threshold. As shown in Figure 5, the black dots represent the first data H to be inferred, and the white dots represent historical data in the scene. The first historical data to be inferred is obtained through neighborhood sampling within the solid-line ellipse. With the corresponding tags Composition of data-label pairs Θ represents the set of data-label pairs in the scene to which the first data H to be inferred belongs. Several data-label pairs In Together with H, they form the first data sequence. With xinit Form the first label sequence.

[0179] In this implementation, a certain amount of first historical data to be inferred can be obtained by sampling the neighborhood of the first data to be inferred. The first historical data to be inferred in the neighborhood usually has a high similarity with the first data to be inferred in the feature space, which can reflect the local structure, trend and pattern of the data. Using these data and the first data to be inferred to construct an input data sequence is helpful for electronic devices to understand the meaning of the first data to be inferred in a specific context or scenario, thereby improving the inference accuracy based on the first data to be inferred.

[0180] For example, another way to obtain the first data sequence and the first label sequence is as follows: Obtain an estimated inference result of the first data to be inferred, where the estimated inference result can be understood as a pseudo-label of the first data to be inferred. The electronic device performs neighborhood sampling on the estimated inference result in the label feature space to obtain a preset number of first target labels. For example, the label feature space can be Euclidean space, spherical space, manifold space, etc. The first target label can be a label whose feature distance from the estimated inference result is less than or equal to a preset threshold. It should be understood that these primary target labels There is corresponding historical data to be inferred. This resulted in several data-label pairs. Several data-label pairs In Together with H, they form the first data sequence. With x init The first tag sequence is formed. The estimated inference result of the first data to be inferred can be obtained by the current electronic device or by other electronic devices. The current electronic device can obtain the estimated inference result from other electronic devices.

[0181] In this implementation, the electronic device can acquire the estimated inference result of the first data to be inferred, and perform neighborhood sampling on the estimated inference result in the label feature space to obtain a first target label with a high similarity to the estimated inference result. Similar labels usually also have first historical data to be inferred that have a certain degree of similarity to the first data to be inferred. Using these data and the first data to be inferred to construct an input data sequence helps the electronic device understand the meaning of the first data to be inferred in a specific context or scenario, thereby improving the inference accuracy based on the first data to be inferred.

[0182] 402: Based on the first data sequence, obtain the first association weight between each data in the first data to be inferred and the first historical data to be inferred.

[0183] In this embodiment, the electronic device can perform dimensional transformation on the first data sequence to obtain a sequence of a specific dimension, and calculate the first association weight between the data in the first data to be inferred and the first historical data to be inferred based on the elements corresponding to each data in the sequence. For example, the electronic device can also use the first data sequence as input to a trained neural network model, and process the first data sequence through the neural network model to obtain the first association weight. For example, the first association weight can be Pearson correlation coefficient, Euclidean distance, mutual information, attention weight, etc.

[0184] 403: Based on the first association weight and the first label sequence, obtain the reasoning result corresponding to the first data to be reasoned.

[0185] In this embodiment, the electronic device can obtain a first feature tensor of the first label sequence based on the first label sequence. For example, it can obtain the reasoning result of the first data to be reasoned by combining the first association weight and the first feature tensor through preprocessing, embedding, linear transformation, etc. For example, the electronic device can concatenate the first association weight and the first feature tensor, use the concatenated tensor as the input of the model, and obtain the reasoning result corresponding to the first data to be reasoned through the model. For example, the electronic device can also extract features from the first label sequence through the neural network model in step 402 to obtain the corresponding first feature tensor. For example, if the neural network model has an attention layer, the first feature tensor can be a value (V) tensor. Based on obtaining the first feature tensor, the electronic device continues to reason through the first association weight and the first feature tensor through the neural network model to obtain the reasoning result of the first data to be reasoned, that is, to obtain the output of the neural network model.

[0186] In this implementation, the first feature tensor can reflect the high-dimensional and complex interactions and temporal patterns between labels in the label sequence. Combining this feature tensor with the first association weight can provide the electronic device with richer and more comprehensive information, which can help the electronic device to understand the relationship between the data in the first data sequence more accurately, thereby improving the reasoning accuracy of the first data to be reasoned.

[0187] As can be seen, in this embodiment, the electronic device can obtain the first historical data to be inferred and the first association weight between each data in the first data sequence based on the first data sequence, and obtain the inference result corresponding to the first data to be inferred by combining the first association weight and the first label sequence corresponding to the first data sequence. Since the first association weight can characterize the causal relationship, correlation or change pattern between each data in the first data to be inferred and the first historical data to be inferred, the electronic device can infer the first data to be inferred based on this association relationship between the data and the labels of the historical data, thereby obtaining the inference result corresponding to the first data to be inferred. This eliminates the need to use data from the scene to which the first data to be inferred belongs to specifically train or fine-tune the AI ​​model for inferring the first data to be inferred, which helps to avoid the various hardware and software overheads brought about by training or fine-tuning the AI ​​model. The inference of the first data to be inferred does not have to wait for the retraining or fine-tuning of the AI ​​model, saving the intermediate training latency. In addition, even if the electronic device obtains the reasoning result corresponding to the first data to be reasoned through the trained neural network model, the neural network model can also obtain the reasoning result corresponding to the first data to be reasoned based on the first association weight and the first label sequence, without the need to fine-tune the neural network model. This can also achieve the effect of reducing overhead and saving latency, and can also realize the cross-scene generalization of the neural network model.

[0188] Please refer to Figure 6, which is a flowchart illustrating another data processing method provided in an embodiment of this application. As shown in Figure 6, the method includes steps 601-604:

[0189] 601: Obtain the first data sequence and the first label sequence. The first data sequence includes the first historical data to be inferred and the first data to be inferred. The labels in the first label sequence correspond one-to-one with the data in the first data sequence in terms of sequence position.

[0190] Step 601 can refer to the relevant description in step 401 above, and can achieve the same or similar beneficial effects.

[0191] 602: Based on the first data sequence, obtain the first association weight between each data in the first data to be inferred and the first historical data to be inferred.

[0192] 603: Obtain the first feature tensor of the first label sequence based on the first label sequence.

[0193] Steps 602 and 603 can be executed sequentially; for example, step 602 can be executed before or after step 603. Alternatively, steps 602 and 603 can be executed in parallel.

[0194] 604: The reasoning result of the first data to be reasoned is obtained based on the first association weight and the first feature tensor.

[0195] Steps 602-604 are executed using a neural network model. In this embodiment, as shown in Figure 7, the neural network model includes N (N≥1) encoder layers. Each of the N encoder layers includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The widths of the first and second modules in the N encoder layers are the same. The first modules in the N encoder layers constitute a data Transformer, and the second modules in the N encoder layers constitute a label Transformer. The first module in the first encoder layer is connected to a first linear layer, the second module in the first encoder layer is connected to a second linear layer, and the second module in the Nth encoder layer is connected to a third linear layer. When N=1, the Nth encoder layer is the same as the first encoder layer; that is, the first encoder layer and the Nth encoder layer are the same encoder layer. Specifically, the query (Q) tensor and key (K) tensor obtained in the first module of each encoder layer are used in the second module of the encoder layer to combine with the V tensor obtained in the second module to obtain the output tensor of the second module. Alternatively, the attention weights obtained in the first module of each encoder layer are used in the second module of the encoder layer to combine with the V tensor obtained in the second module to obtain the output tensor of the second module. That is, the first module of each encoder layer is connected to the second module of the encoder layer through the Q tensor and the K tensor, or the first module of each encoder layer is connected to the second module of the encoder layer through the attention weights.

[0196] For example, please continue to refer to Figure 7, the electronic device will display the first data sequence (such as...) ) as input to the first Linear layer, and the first label sequence (such as As input to the second Linear layer, the corresponding data is consistent with the position of its label in its respective sequence.

[0197] When N > 1, the first Linear layer performs a dimensionality transformation on the first data sequence to obtain a tensor corresponding to the first data sequence. This tensor serves as the input to the first module in the first encoder layer. The electronic device processes this tensor through the first module in the first encoder layer to obtain the first Q tensor, the first K tensor, and the first V2 tensor. The second Linear layer performs a dimensionality transformation on the first label sequence to obtain a tensor corresponding to the first label sequence. This tensor is added to the position embedding of the first label sequence to obtain the input to the second module in the first encoder layer. The electronic device processes this input through the second module in the first encoder layer to obtain the first V1 tensor. Based on the first Q tensor and the first K tensor, the electronic device obtains the first attention weight. The second module in the first encoder layer processes the first attention weight and the first V1 tensor to obtain the output tensor of the second module in the first encoder layer. The electronic device processes the first Q tensor, the first K tensor, and the first V2 tensor through the first module in the first encoder layer to obtain the output tensor of the first module in the first encoder layer. For the i-th encoder layer out of N encoder layers, the output tensor of the first module in the (i-1)-th encoder layer is the input of the first module in the i-th encoder layer, and the output tensor of the second module in the (i-1)-th encoder layer is the input of the second module in the i-th encoder layer, where 1 < i ≤ N. The electronic device processes the output tensor of the first module in the (i-1)-th encoder layer through the first module in the i-th encoder layer to obtain the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor. It then further processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor through the first module to obtain the output tensor of the first module in the i-th encoder layer. The output tensor of the first module in the N-th encoder layer is not used as the output of the neural network model. The electronic device obtains the i-th attention weight based on the i-th Q tensor and the i-th K tensor. It then processes the output tensor of the second module in the (i-1)-th encoder layer through the second module in the i-th encoder layer to obtain the i-th V1 tensor. The electronic device further processes the i-th attention weight and the i-th V1 tensor through the second module in the i-th encoder layer to obtain the output tensor of the second module in the i-th encoder layer. Based on the above description, having obtained the output tensor of the second module in the N-th encoder layer, the electronic device, through inference using the neural network model, obtains N attention weights and N V1 tensors. The first association weight includes these N attention weights, and the first feature tensor includes these N V1 tensors.

[0198] It should be noted that position embedding can be implemented in the label Transformer, the data Transformer, or both the label Transformer and the data Transformer.

[0199] In this implementation, the electronic device processes the first label sequence through a linear layer to obtain the input tensor of the second module in the first encoder layer. This second module then processes the input tensor to obtain the first V1 tensor. Based on this first V1 tensor and the first attention weight, the electronic device obtains the output tensor of the second module in the first encoder layer. For subsequent encoder layers, the second module, based on the output tensor of the second module in the previous encoder layer, obtains the V1 tensor extracted by the second module in the current encoder layer. Based on this, the neural network model with N encoder layers will extract N V1 tensors. These N V1 tensors contain relevant information from the first label sequence. Combining this information with the corresponding attention weights for attention calculation provides additional guidance to the neural network model, helping it focus on labels related to the first data to be inferred, thereby improving the model's inference performance on that data.

[0200] Based on the connection between the first and second modules in each encoder layer, the electronic device can obtain the attention weights generated in each encoder layer. It then processes these attention weights and the V1 tensor generated by the second module in the current encoder layer to obtain the output tensor of the second module in the current encoder layer. Having obtained the output tensor of the second module in the Nth encoder layer, the electronic device processes this output tensor through a third Linear layer to obtain the inference result for the first data to be inferred. For each encoder layer, the second module obtains attention weights through its connection with the first module and combines the corresponding V1 tensor with these attention weights. This introduces label information to guide the neural network model inference to obtain the output tensor of the second module in the current encoder layer, which helps improve the neural network model's understanding of the first data to be inferred and enhances the model's interpretability.

[0201] In addition, the first module in each encoder layer focuses on the attention mechanism processing at the data level. The Q tensor, K tensor and V2 tensor generated by the first module in the current encoder layer are used to generate the output tensor of the first module, so as to provide input for the first module in the next encoder layer, thereby enabling the next encoder layer to generate attention weights based on the input.

[0202] For example, the structure of any one of the N encoder layers can be as shown in Figure 8. The first and second modules in the encoder layer both include a masked multi-head attention layer and a feed-forward network layer. Both the masked multi-head attention layer and the feed-forward network layer are preceded by a normalization layer.

[0203] In the first module of the encoder layer, a normalization layer preceding the masked multi-head attention layer normalizes the input tensor. The normalized tensor generates Q-tensor, K-tensor, and V2 tensors based on the corresponding Q-weight matrix, K-weight matrix, and V-weight matrix, respectively. These Q-tensor, K-tensor, and V2 tensors serve as inputs to the masked multi-head attention layer. The masked multi-head attention layer generates attention weights based on the K-tensor and Q-tensor. These attention weights, combined with the V2 tensor, generate the output of the masked multi-head attention layer. The masked multi-head attention layer can mask the features corresponding to the first data to be inferred when calculating attention. The output of the masked multi-head attention layer is added to the input of the first module to obtain the input of the normalization layer preceding the feedforward network layer. The output of this normalization layer serves as the input of the feedforward network layer, which processes the data to obtain its corresponding output. Finally, the input of the normalization layer preceding the feedforward network layer is added to the output of the feedforward network layer to obtain the output tensor of the first module of the encoder layer.

[0204] In the second module of the encoder layer, a normalization layer preceding the masked multi-head attention layer normalizes the input tensor. The normalized tensor generates a V1 tensor based on the corresponding V weight matrix. The Q tensor, K tensor generated in the first module, and this V1 tensor serve as the input to the masked multi-head attention layer. The masked multi-head attention layer generates attention weights based on the K and Q tensors. These attention weights, combined with the V1 tensor, generate the output of the masked multi-head attention layer. Specifically, the masked multi-head attention layer can mask the features corresponding to the initial labels of the first data to be inferred when calculating attention. The output of the masked multi-head attention layer is added to the input of the second module to obtain the input of the normalization layer preceding the feedforward network layer. The output of this normalization layer serves as the input of the feedforward network layer, which processes the data to obtain its corresponding output. The input of the normalization layer preceding the feedforward network layer is then added to the output of the feedforward network layer to obtain the output tensor of the second module of the encoder layer.

[0205] In other words, for the first encoder layer, the electronic device can process the first Q tensor and the first K tensor through its second module to obtain the first attention weight. For the i-th encoder layer, the electronic device can process the i-th Q tensor and the i-th K tensor through its second module to obtain the i-th attention weight.

[0206] For example, the structure of any one of the N encoder layers can also be as shown in Figure 9. The Q tensor, K tensor, and V2 tensor generated by the first module of the encoder layer serve as inputs to the mask-based multi-head attention layer. The mask-based multi-head attention layer generates attention weights based on the K tensor and Q tensor. These attention weights are combined with the V2 tensor to generate the output of the mask-based multi-head attention layer. The first module performs forward inference based on this output. After generating the attention weights, the mask-based multi-head attention layer in the first module transmits the attention weights to the second module. The V1 tensor generated in the second module and the attention weights serve as inputs to the mask-based multi-head attention layer. The output of this layer is obtained through the processing of the mask-based multi-head attention layer, and the second module performs forward inference based on this output. The difference between the encoder layer structure shown in Figure 9 and the encoder layer structure shown in Figure 8 is that the first module and the second module are connected through attention weights, instead of transmitting the Q tensor and K tensor to the second module. Other processing can be described in the corresponding descriptions in the structure shown in Figure 8.

[0207] In other words, for the first encoder layer, the electronic device can process the first Q tensor and the first K tensor through its first module to obtain the first attention weight. For the i-th encoder layer, the electronic device can process the i-th Q tensor and the i-th K tensor through its first module to obtain the i-th attention weight.

[0208] In this implementation, attention weights in each encoder layer can be obtained either by the first module based on the Q-tensor and K-tensor generated by the first module, or by the second module based on the Q-tensor and K-tensor generated by the first module. The generation of attention weights is relatively flexible due to the different connection methods between the first and second modules.

[0209] When N=1, the first Linear layer performs a dimensionality transformation on the first data sequence to obtain a tensor corresponding to the first data sequence. This tensor serves as the input to the first module in the encoder layer. The electronic device processes this tensor through the first module in the encoder layer to obtain the Q tensor, K tensor, and V2 tensor. The second Linear layer performs a dimensionality transformation on the first label sequence to obtain a tensor corresponding to the first label sequence. This tensor is added to the position embedding of the first label sequence to obtain the input to the second module in the encoder layer. The electronic device processes this input through the second module in the encoder layer to obtain the V1 tensor. Based on the Q tensor and K tensor, the electronic device obtains the attention weights. The second module in the encoder layer processes the attention weights and the V1 tensor to obtain the output tensor of the second module in the encoder layer. The electronic device processes the Q tensor, K tensor, and V2 tensor through the first module in the encoder layer to obtain the output tensor of the first module in the encoder layer. The output tensor of the first module in the encoder layer is not used as the output of the neural network model. The electronic device processes the output tensor of the second module in the encoder layer through the third Linear layer to obtain the inference result of the first data to be inferred. The beneficial effects of the implementation with N=1 can be compared with the beneficial effects of the implementation with N>1. The structure of the encoder layer can be seen in the description in Figure 8 or Figure 9, and will not be repeated here.

[0210] Please refer to Figure 10, which is a flowchart illustrating a training method for a neural network model provided in an embodiment of this application. This method can be executed by an electronic device, which may be the electronic device shown in the embodiments of Figures 4 and 6, or it may not be the electronic device shown in the embodiments of Figures 4 and 6. When it is not the electronic device shown in the embodiments of Figures 4 and 6, for ease of distinction, the electronic device shown in the embodiments of Figures 4 and 6 may be a first electronic device, while the electronic device here may be a second electronic device. As shown in Figure 10, the method includes steps 1001-1003:

[0211] 1001: Obtain the second data sequence and the second label sequence. The second data sequence includes the second historical data to be inferred and the second data to be inferred. The labels in the second label sequence correspond one-to-one with the data in the second data sequence in terms of sequence position.

[0212] In this embodiment of the application, the second historical data to be inferred and the second data to be inferred can be training data from several data-label pairs extracted from the current scene training set. Each data in the training data has a corresponding label. For example, in a scenario where CSI is used for user positioning, each CSI in the second data sequence has corresponding location information.

[0213] For example, one implementation of obtaining the second data sequence and the second label sequence is as follows: neighborhood sampling is performed on the second data to be inferred (the second data to be inferred can be any training data in the training set) in the data feature space to obtain a preset number of second historical data to be inferred; this preset number of second historical data to be inferred is combined with the second data to be inferred to form the second data sequence; and the labels of the second historical data to be inferred are combined with the initial labels of the second data to be inferred to form the second label sequence. The specific implementation of obtaining the second historical data to be inferred through neighborhood sampling can refer to the relevant description of obtaining the first historical data to be inferred through neighborhood sampling of the first data to be inferred.

[0214] In this implementation, a certain number of second historical data to be inferred can be obtained by sampling the neighborhood of the second data to be inferred. The second historical data to be inferred in the neighborhood usually has a high similarity with the second data to be inferred in the feature space, which can reflect the local structure, trend and pattern of the data. Using these data and the second data to be inferred to construct the input data sequence is helpful for electronic devices to understand the meaning of the second data to be inferred in a specific context or scenario, thereby improving the inference accuracy of the second data to be inferred.

[0215] For example, another implementation of obtaining the second data sequence and the second label sequence is as follows: neighborhood sampling is performed on the labels of the second data to be inferred in the label feature space to obtain a preset number of second target labels; the historical data to be inferred corresponding to the second target labels is used as the second historical data to be inferred; the preset number of second historical data to be inferred and the second data to be inferred are combined to form the second data sequence; and the second target labels of the second historical data to be inferred and the initial labels of the second data to be inferred are combined to form the second label sequence. The specific implementation of obtaining the second target label through neighborhood sampling can refer to the relevant description of obtaining the first target label through neighborhood sampling of the estimated inference result of the first data to be inferred.

[0216] In this implementation, the electronic device can perform neighborhood sampling on the labels of the second data to be inferred in the label feature space to obtain a second target label with high similarity to the first label. Similar labels usually also have second historical data to be inferred that have a certain degree of similarity to the second data to be inferred. Using these data and the second data to be inferred to construct an input data sequence helps the electronic device understand the meaning of the second data to be inferred in a specific context or scenario, thereby improving the inference accuracy of the second data to be inferred.

[0217] 1002: Perform the following processing through a neural network: obtain the second association weight between each data in the second data to be inferred and the second historical data to be inferred based on the second data sequence, and obtain the inference result of the second data to be inferred based on the second association weight and the second label sequence.

[0218] For example, an electronic device can obtain a second feature tensor of the second label sequence based on the second label sequence, and then combine the second association weights and the second feature tensor to obtain the reasoning result of the second data to be reasoned. For example, the second feature tensor can be obtained through preprocessing, embedding, linear transformation, etc. The electronic device can concatenate the second association weights with the second feature tensor, and use the concatenated tensor as the input of the inferencer to obtain the reasoning result of the second data to be reasoned.

[0219] For example, the electronic device can also extract features from the second label sequence using a neural network to obtain the corresponding second feature tensor. For instance, if the neural network has an attention layer, the second feature tensor can be a V tensor. Based on the obtained second feature tensor, the electronic device continues to reason about the second association weights and the second feature tensor using the neural network to obtain the reasoning result of the second data to be reasoned, i.e., the output of the neural network.

[0220] In this implementation, the second feature tensor can reflect the high-dimensional and complex interactions and temporal patterns between labels in the label sequence. Combining this feature tensor with the second association weight can provide the electronic device with richer and more comprehensive information, which can help the electronic device to more accurately understand the association between the data in the second data sequence, thereby improving the reasoning accuracy of the second data to be reasoned.

[0221] In this embodiment, the neural network structure can be seen in Figure 7, including N (N≥1) encoder layers. Each encoder layer includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The first module in the first encoder layer is connected to a first linear layer, the second module in the first encoder layer is connected to a second linear layer, and the second module in the Nth encoder layer is connected to a third linear layer. When N=1, the Nth encoder layer is the same as the first encoder layer; that is, the first encoder layer and the Nth encoder layer are the same encoder layer.

[0222] Specifically, the electronic device uses the second data sequence as input to the first Linear layer and the second tag sequence as input to the second Linear layer, with the corresponding data and tags maintaining the position of their respective tags within their sequences.

[0223] When N > 1, the first Linear layer performs a dimensionality transformation on the second data sequence to obtain a tensor corresponding to the second data sequence. This tensor serves as the input to the first module in the first encoder layer. The electronic device processes this tensor through the first module in the first encoder layer to obtain the first Q tensor, the first K tensor, and the first V2 tensor. The second Linear layer performs a dimensionality transformation on the second label sequence to obtain a tensor corresponding to the second label sequence. This tensor is added to the position embedding of the second label sequence to obtain the input to the second module in the first encoder layer. The electronic device processes this input through the second module in the first encoder layer to obtain the first V1 tensor. Based on the first Q tensor and the first K tensor, the electronic device obtains the first attention weight. The first attention weight and the first V1 tensor are then processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The electronic device processes the first Q tensor, the first K tensor, and the first V2 tensor through the first module in the first encoder layer to obtain the output tensor of the first module in the first encoder layer. For the i-th encoder layer out of N encoder layers, the output tensor of the first module in the (i-1)-th encoder layer is the input of the first module in the i-th encoder layer, and the output tensor of the second module in the (i-1)-th encoder layer is the input of the second module in the i-th encoder layer, where 1 < i ≤ N. The electronic device processes the output tensor of the first module in the (i-1)-th encoder layer through the first module in the i-th encoder layer to obtain the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor. It then further processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor through the first module to obtain the output tensor of the first module in the i-th encoder layer. The output tensor of the first module in the N-th encoder layer is not used as the output of the neural network. The electronic device obtains the i-th attention weight based on the i-th Q-tensor and the i-th K-tensor. It then processes the output tensor of the second module in the (i-1)-th encoder layer through the second module in the i-th encoder layer to obtain the i-th V1 tensor. The electronic device further processes the i-th attention weight and the i-th V1 tensor through the second module in the i-th encoder layer to obtain the output tensor of the second module in the i-th encoder layer. Based on the above description, having obtained the output tensor of the second module in the N-th encoder layer, the electronic device, through neural network inference, obtains N attention weights and N V1 tensors. The second association weight includes these N attention weights, and the second feature tensor includes these N V1 tensors.

[0224] It should be noted that position embedding can be implemented in the label Transformer, the data Transformer, or both the label Transformer and the data Transformer.

[0225] In this implementation, the electronic device processes the second label sequence through a linear layer to obtain the input tensor of the second module in the first encoder layer. This second module then processes the input tensor to obtain the first V1 tensor. Based on this first V1 tensor and the first attention weight, the electronic device obtains the output tensor of the second module in the first encoder layer. For subsequent encoder layers, the second module, based on the output tensor of the second module in the previous encoder layer, obtains the V1 tensor extracted by the second module in the current encoder layer. Thus, a neural network with N encoder layers extracts N V1 tensors. These N V1 tensors contain relevant information from the second label sequence. Combining this information with the corresponding attention weights for attention calculation provides additional guidance to the neural network, helping it focus on labels related to the second data to be inferred, thereby improving the neural network's inference performance on the second data.

[0226] Based on the connection between the first and second modules in each encoder layer, the electronic device can obtain the attention weights generated in each encoder layer. It then processes these attention weights and the V1 tensor generated by the second module in the current encoder layer to obtain the output tensor of the second module in the current encoder layer. Having obtained the output tensor of the second module in the Nth encoder layer, the electronic device processes this output tensor through a third linear layer to obtain the inference result for the second data to be inferred. For each encoder layer, the second module obtains attention weights through its connection with the first module and combines the corresponding V1 tensor with these attention weights. This introduces label information to guide the neural network inference, resulting in the output tensor of the second module in the current encoder layer. This improves the neural network's understanding of the second data to be inferred and enhances the interpretability of the neural network.

[0227] In addition, the first module in each encoder layer focuses on the attention mechanism processing at the data level. The Q tensor, K tensor and V2 tensor generated by the first module in the current encoder layer are used to generate the output tensor of the first module, so as to provide input for the first module in the next encoder layer, thereby enabling the next encoder layer to generate attention weights based on the input.

[0228] For example, the structure of any one of the N encoder layers can be as shown in Figure 8 or Figure 9. Specifically, for the first encoder layer, the electronic device can process the first Q tensor and the first K tensor through its second module to obtain the first attention weight; for the i-th encoder layer, the electronic device can process the i-th Q tensor and the i-th K tensor through its second module to obtain the i-th attention weight. Alternatively, for the first encoder layer, the electronic device can process the first Q tensor and the first K tensor through its first module to obtain the first attention weight; for the i-th encoder layer, the electronic device can process the i-th Q tensor and the i-th K tensor through its first module to obtain the i-th attention weight.

[0229] In this implementation, attention weights in each encoder layer can be obtained either by the first module based on the Q-tensor and K-tensor generated by the first module, or by the second module based on the Q-tensor and K-tensor generated by the first module. The generation of attention weights is relatively flexible due to the different connection methods between the first and second modules.

[0230] When N=1, the first Linear layer performs a dimensionality transformation on the second data sequence to obtain a tensor corresponding to the second data sequence. This tensor serves as the input to the first module in the encoder layer. The electronic device processes this tensor through the first module in the encoder layer to obtain the Q tensor, K tensor, and V2 tensor. The second Linear layer performs a dimensionality transformation on the second label sequence to obtain a tensor corresponding to the second label sequence. This tensor is added to the position embedding of the second label sequence to obtain the input to the second module in the encoder layer. The electronic device processes this input through the second module in the encoder layer to obtain the V1 tensor. Based on the Q tensor and K tensor, the electronic device obtains the attention weights. The second module in the encoder layer processes the attention weights and the V1 tensor to obtain the output tensor of the second module in the encoder layer. The electronic device processes the Q tensor, K tensor, and V2 tensor through the first module in the encoder layer to obtain the output tensor of the first module in the encoder layer. The output tensor of the first module in the encoder layer is not used as the output of the neural network. The electronic device processes the output tensor of the second module in the encoder layer through the third Linear layer to obtain the inference result of the second data to be inferred. The beneficial effects of the implementation with N=1 can be compared with the beneficial effects of the implementation with N>1. The structure of the encoder layer can be seen in the description in Figure 8 or Figure 9, and will not be repeated here.

[0231] 1003: Adjust the parameters of the neural network based on the inference results and labels of the second inference data to obtain the neural network model.

[0232] In this embodiment, the loss of the second data to be inferred can be obtained based on the inference result and the label of the second data to be inferred. Based on this loss, the parameters of the neural network can be adjusted by the backpropagation algorithm. After multiple rounds of iteration of the training data sequence and label sequence, a trained neural network model can be obtained. This neural network model can be used to perform inference of the first data to be inferred in the embodiment shown in Figure 4 or Figure 6.

[0233] As can be seen, in this embodiment, the electronic device can obtain the second correlation weight between the second historical data to be reasoned and the data in the second data to be reasoned based on the second data sequence, and combine the second correlation weight with the second label sequence corresponding to the second data sequence to obtain the reasoning result of the second data to be reasoned. Since the second correlation weight can characterize the causal relationship, correlation, or change pattern between the data in the second data to be reasoned and the data in the second historical data to be reasoned, the electronic device can reason about the second data to be reasoned based on this correlation relationship between the data and the labels of the historical data, thereby obtaining the reasoning result of the second data to be reasoned. Based on the reasoning result of the second data to be reasoned and its labels, the electronic device can adjust the parameters of the neural network to obtain a neural network model that can be directly applied to various scenarios for data reasoning. Since the neural network learns to reason about the current data based on the correlation relationship between the data and the labels of the historical data during the training phase, the neural network model does not need to be specially trained or fine-tuned using the data of that scenario when applied to different scenarios. This helps to avoid the various hardware and software overheads brought about by training or fine-tuning the neural network model, and can be directly used for reasoning of scenario data, which helps to save intermediate training or fine-tuning latency and realizes the cross-scenario generalization of the neural network model.

[0234] The methods of the embodiments of this application have been described above. The data processing apparatus and the neural network model training apparatus of the embodiments of this application are provided below.

[0235] Please refer to Figure 11, which is a schematic diagram of a data processing device provided in an embodiment of this application. As shown in Figure 11, the device includes at least a first acquisition unit 1101 and a first processing unit 1102, wherein:

[0236] The first acquisition unit 1101 is used to acquire a first data sequence and a first tag sequence. The first data sequence includes first historical data to be inferred and first data to be inferred. The tags in the first tag sequence correspond one-to-one with the data in the first data sequence in terms of sequence position.

[0237] The first processing unit 1102 is used to obtain the first association weight between each data in the first data to be reasoned and the first historical data to be reasoned based on the first data sequence; and to obtain the reasoning result of the first data to be reasoned based on the first association weight and the first label sequence.

[0238] In one possible implementation, the first processing unit 1102 is further configured to: perform neighborhood sampling on the first data to be inferred in the data feature space to obtain the first historical data to be inferred.

[0239] In one possible implementation, the first acquisition unit 1101 is further configured to: acquire the estimated reasoning result of the first data to be reasoned; the first processing unit 1102 is further configured to: perform neighborhood sampling on the estimated reasoning result in the label feature space to obtain the first target label, and the first historical data to be reasoned is the historical data to be reasoned corresponding to the first target label.

[0240] In one possible implementation, in obtaining the reasoning result of the first data to be reasoned based on the first association weight and the first label sequence, the first processing unit 1102 is specifically used for:

[0241] The first feature tensor of the first label sequence is obtained based on the first label sequence;

[0242] The reasoning result of the first data to be reasoned is obtained based on the first association weight and the first feature tensor.

[0243] In one possible implementation, the reasoning result of the first data to be reasoned based on the first association weight and the first label sequence is obtained by a neural network model. The neural network model includes N encoder layers, and each of the N encoder layers includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism.

[0244] The first feature tensor includes N value V1 tensors. Regarding obtaining the first feature tensor of the first label sequence based on the first label sequence, the first processing unit 1102 is specifically used for:

[0245] The tensor obtained based on the first label sequence is processed by the second module in the first encoder layer of N encoder layers to obtain the first V1 tensor among N V1 tensors;

[0246] The output tensor of the second module in the (i-1)th encoder layer of N encoder layers is processed by the second module in the i-th encoder layer to obtain the i-th V1 tensor among N V1 tensors, where 1 < i ≤ N.

[0247] In one possible implementation, the first association weights include N attention weights. Regarding obtaining the reasoning result of the first data to be reasoned based on the first association weights and the first feature tensor, the first processing unit 1102 is specifically used for:

[0248] The first attention weight and the first V1 tensor out of N attention weights are processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The first attention weight is obtained based on the first query Q tensor and the first key K tensor. The first Q tensor and the first K tensor are obtained by the first module in the first encoder layer processing the tensor obtained based on the first data sequence.

[0249] The second module in the i-th encoder layer processes the i-th attention weight and the i-th V1 tensor among the N attention weights to obtain the output tensor of the second module in the i-th encoder layer. The i-th attention weight is obtained based on the i-th Q tensor and the i-th K tensor. The i-th Q tensor and the i-th K tensor are obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0250] When i equals N, the reasoning result of the first data to be reasoned is obtained based on the output tensor of the second module in the Nth encoder layer.

[0251] In one possible implementation, the first processing unit 1102 is further configured to:

[0252] The first module in the first encoder layer processes the first Q tensor, the first K tensor, and the first V2 tensor to obtain the output tensor of the first module in the first encoder layer. The first V2 tensor is obtained by the first module in the first encoder layer processing the tensor based on the first data sequence.

[0253] The first module in the i-th encoder layer processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor to obtain the output tensor of the first module in the i-th encoder layer. The i-th V2 tensor is obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0254] In one possible implementation, the first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer.

[0255] The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the second module in the i-th encoder layer.

[0256] It should be noted that the implementation of each unit described in FIG11 can also be described with reference to the corresponding descriptions of the embodiments shown in FIG4 to FIG9. Furthermore, the beneficial effects of the data processing apparatus described in FIG11 can be described with reference to the corresponding descriptions of the embodiments shown in FIG4 to FIG9, and will not be repeated here.

[0257] The above embodiments use a neural network model as an example to illustrate the technical solution. The above solution can also be applied to other models, and this application does not impose any limitations.

[0258] Please refer to Figure 12, which is a schematic diagram of a training device for a neural network model provided in an embodiment of this application. As shown in Figure 12, the device includes at least a second acquisition unit 1201 and a second processing unit 1202, wherein:

[0259] The second acquisition unit 1201 is used to acquire a second data sequence and a second tag sequence. The second data sequence includes second historical data to be reasoned and second data to be reasoned. The tags in the second tag sequence correspond one-to-one with the data in the second data sequence in terms of sequence position.

[0260] The second processing unit 1202 is used to perform the following processing through a neural network: obtaining the second association weight between each data in the second data to be inferred and the second historical data to be inferred based on the second data sequence, and obtaining the inference result of the second data to be inferred based on the second association weight and the second label sequence;

[0261] The second processing unit 1202 is also used to adjust the parameters of the neural network based on the reasoning result of the second data to be reasoned and the label of the second data to be reasoned, so as to obtain a neural network model.

[0262] In one possible implementation, the second processing unit 1202 is further configured to: perform neighborhood sampling on the second data to be inferred in the data feature space to obtain the second historical data to be inferred.

[0263] In one possible implementation, the second processing unit 1202 is further configured to: perform neighborhood sampling on the labels of the second data to be inferred in the label feature space to obtain the second target label, and the second historical data to be inferred is the historical data to be inferred corresponding to the second target label.

[0264] In one possible implementation, the second processing unit 1202 is specifically used for: obtaining the reasoning result of the second data to be reasoned based on the second association weight and the second label sequence;

[0265] The second feature tensor of the second label sequence is obtained based on the second label sequence;

[0266] The reasoning result of the second data to be reasoned is obtained based on the second association weight and the second feature tensor.

[0267] In one possible implementation, the neural network includes N encoder layers, each of which includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism.

[0268] The second feature tensor includes N V1 tensors. Regarding obtaining the second feature tensor of the second label sequence based on the second label sequence, the second processing unit 1202 is specifically used for:

[0269] The tensor obtained based on the second label sequence is processed by the second module in the first encoder layer of N encoder layers to obtain the first V1 tensor among N V1 tensors;

[0270] The output tensor of the second module in the (i-1)th encoder layer of N encoder layers is processed by the second module in the i-th encoder layer to obtain the i-th V1 tensor among N V1 tensors, where 1 < i ≤ N.

[0271] In one possible implementation, the second association weights include N attention weights. Regarding the reasoning result obtained from the second inference data based on the second association weights and the second feature tensor, the second processing unit 1202 is specifically used for:

[0272] The first attention weight and the first V1 tensor out of N attention weights are processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The first attention weight is obtained based on the first query Q tensor and the first key K tensor. The first Q tensor and the first K tensor are obtained by the first module in the first encoder layer processing the tensor obtained based on the second data sequence.

[0273] The second module in the i-th encoder layer processes the i-th attention weight and the i-th V1 tensor among the N attention weights to obtain the output tensor of the second module in the i-th encoder layer. The i-th attention weight is obtained based on the i-th Q tensor and the i-th K tensor. The i-th Q tensor and the i-th K tensor are obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0274] When i equals N, the reasoning result of the second data to be reasoned is obtained based on the output tensor of the second module in the Nth encoder layer.

[0275] In one possible implementation, the second processing unit 1202 is further configured to:

[0276] The first module in the first encoder layer processes the first Q tensor, the first K tensor, and the first V2 tensor to obtain the output tensor of the first module in the first encoder layer. The first V2 tensor is obtained by the first module in the first encoder layer processing the tensor based on the second data sequence.

[0277] The first module in the i-th encoder layer processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor to obtain the output tensor of the first module in the i-th encoder layer. The i-th V2 tensor is obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

[0278] In one possible implementation, the first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer.

[0279] The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor through the second module in the i-th encoder layer.

[0280] It should be noted that the implementation of each unit described in FIG12 can also correspond to the description of the embodiment shown in FIG10. Furthermore, the beneficial effects of the training device for the neural network model described in FIG12 can be described in the corresponding description of the embodiment shown in FIG10, and will not be repeated here.

[0281] Based on the description of the above method and device embodiments, this application also provides an electronic device. Please refer to FIG13, which is a schematic diagram of the structure of an electronic device provided in this application embodiment. The electronic device includes at least one processor 1301. Optionally, the electronic device may further include an interface circuit 1302 (shown as dashed lines in the figure), with the processor 1301 and the interface circuit 1302 coupled to each other. It is understood that the interface circuit 1302 can be a transceiver or an input / output interface. Optionally, the electronic device may further include at least one memory 1303 (shown as dashed lines in the figure), which is used to store instructions (such as one or more computer programs) executed by at least one processor 1301, or to store input data required for at least one processor 1301 to execute instructions, or to store data generated after at least one processor 1301 executes instructions. This electronic device can be used in related steps of a data processing method. The at least one processor 1301 in this electronic device is used to read the computer program code stored in the at least one memory 1303 and execute the method of any one of the embodiments shown in FIG4 to FIG10.

[0282] At least one memory 1303 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM).

[0283] At least one processor 1301 may be one or more central processing units (CPUs). If the processor 1301 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.

[0284] For example, when the electronic device is used to implement the methods shown in Figures 4 to 9, at least one processor 1301 in the electronic device can be used to read one or more programs stored in at least one memory 1303 and perform the following operations:

[0285] Obtain a first data sequence and a first label sequence. The first data sequence includes first historical data to be inferred and first data to be inferred. The labels in the first label sequence correspond one-to-one with the data in the first data sequence in terms of sequence position.

[0286] The first association weights between each data point in the first data to be reasoned and the first historical data to be reasoned are obtained based on the first data sequence.

[0287] The reasoning result of the first data to be reasoned is obtained based on the first association weight and the first label sequence.

[0288] For example, when the electronic device is used to implement the method shown in FIG10, at least one processor 1301 in the electronic device can be used to read one or more programs stored in at least one memory 1303 and perform the following operations:

[0289] Obtain the second data sequence and the second label sequence. The second data sequence includes the second historical data to be reasoned and the second data to be reasoned. The labels in the second label sequence correspond one-to-one with the data in the second data sequence in terms of sequence position.

[0290] The following processing is performed through a neural network: the second association weight between each data in the second data to be inferred and the second historical data to be inferred is obtained based on the second data sequence, and the inference result of the second data to be inferred is obtained based on the second association weight and the second label sequence;

[0291] The parameters of the neural network are adjusted based on the inference results and labels of the second inference data to obtain the neural network model.

[0292] It should be noted that the implementation of each operation can also correspond to the description of the method in any of the embodiments shown in Figures 4 to 10.

[0293] It should be noted that although the electronic device shown in FIG13 only illustrates at least one processor 1301, interface circuit 1302, and at least one memory 1303, those skilled in the art should understand that in specific implementations, the electronic device may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the electronic device may also include hardware devices for implementing other additional functions. Moreover, those skilled in the art should understand that the electronic device may only include the devices necessary for implementing the embodiments of this application, and not necessarily all the devices shown in FIG13.

[0294] This application also provides a chip, including: a processor, configured to retrieve and run a computer program from a memory, causing a device or apparatus on which the chip is mounted to perform the method described in any of the embodiments shown in Figures 4 to 10 above. This chip may be a chip in an electronic device.

[0295] This application also provides a computer-readable storage medium (memory) storing a computer program that, when executed, implements the method described in any of the embodiments shown in Figures 4 to 10. It is understood that the computer-readable storage medium here can include both built-in storage media within a device and extended storage media supported by the device. The computer-readable storage medium provides storage space containing the device's operating system. Furthermore, one or more computer programs suitable for loading and execution by the device's processor are also stored in this storage space. It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor.

[0296] This application also provides a computer program product, which includes computer program code. When the computer program code is run by a device or equipment, the method flow described in any one of the embodiments in Figures 4 to 10 is implemented.

[0297] Please refer to Figure 14, which is a schematic diagram of a baseband hardware provided in an embodiment of this application. As shown in Figure 14, the baseband can be implemented using a processing system including one or more processors. The processor can include a microprocessor, microcontroller, CPU, GPU, or other general-purpose processors, digital signal processors (DSPs), ASICs, FPGAs, programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to various functions. That is, the processor used in the baseband can be used to implement any one or more of the processes described below. It should be understood that the electronic device shown in Figure 13 can be the baseband shown in Figure 14.

[0298] Processing systems can be implemented using a bus architecture, typically represented by a bus. A bus can include any number of interconnect buses and bridges, depending on the specific application and overall design constraints of the processing system. The bus couples various circuits together, including one or more processors (typically represented by a processor), memory, and computer-readable media (typically represented by a computer-readable storage medium). The bus can also link various other circuits, such as timing sources, peripherals, voltage regulators, and power management circuits, which are well-known in the art and will not be described further here. The bus interface provides the interface between the bus and transceivers, as well as between the bus and the interface.

[0299] A transceiver provides a communication interface or means for communicating with various other devices via a wireless transmission medium. The transceiver may be coupled to an antenna array, and the transceiver and antenna array may be used together for communication with a corresponding network type. At least one interface (e.g., a network interface and / or a user interface) provides a communication interface or means for communication via an internal bus or via an external transmission medium.

[0300] The processor is responsible for managing the bus and general processing, including executing software stored on a computer-readable storage medium. When the processor executes the software, it causes the processing system to perform the various functions described below for any particular device. The functions that can be implemented by the processor, memory, and computer-readable medium can include: encoding, decoding, rate matching, rate dematching, scrambling, descrambling, modulation, demodulation, layer mapping, fast Fourier transform (FFT), inverse fast Fourier transform (IFFT), inverse discrete Fourier transform (IDFT), precoding, RE mapping, channel equalization, deRE mapping, digital beamforming (BF), and so on.

[0301] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0302] The processor in this embodiment has signal processing capabilities and can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), neural network processing units (NPUs), artificial intelligence processors, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor can be a microprocessor, any conventional processor, or one or more integrated circuits used to control the execution of a program for controlling the method provided in any of the above embodiments. Some or all steps of the communication method in this embodiment can be implemented by a GPU or NPU, or by a GPU or NPU in conjunction with other processors.

[0303] It should also be understood that the memory mentioned in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be ROM, Programmable Read-Only Memory (PROM), EPROM, Electrically Erasable Programmable Read-Only Memory (EEPROM), or flash memory. Volatile memory can be RAM, which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Synchlink Dynamic Random Access Memory (SLDRAM), and Direct Rambus RAM (DR RAM).

[0304] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) is integrated into the processor.

[0305] It should be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0306] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0307] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely exemplary. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0308] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0309] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0310] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. In the textual description of this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0311] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0312] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0313] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A data processing method, characterized by, The method includes: Obtain a first data sequence and a first tag sequence. The first data sequence includes first historical data to be inferred and first data to be inferred. The tags in the first tag sequence correspond one-to-one with the data in the first data sequence in terms of sequence position. Based on the first data sequence, the first association weight between each data in the first data to be reasoned and the first historical data to be reasoned is obtained; The reasoning result of the first data to be reasoned is obtained based on the first association weight and the first label sequence.

2. The method according to claim 1, characterized in that, The method further includes: The first historical data to be inferred is obtained by sampling the neighborhood of the first data to be inferred in the data feature space.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the estimated reasoning result of the first data to be reasoned; The estimated inference result is sampled in the label feature space to obtain the first target label, and the first historical inference data is the historical inference data corresponding to the first target label.

4. The method according to any one of claims 1-3, characterized in that, The reasoning result obtained based on the first association weight and the first label sequence for the first data to be reasoned includes: Based on the first label sequence, a first feature tensor of the first label sequence is obtained; The reasoning result of the first data to be reasoned is obtained based on the first association weight and the first feature tensor.

5. The method according to claim 4, characterized in that, The reasoning result of the first data to be reasoned based on the first association weight and the first label sequence is obtained by a neural network model. The neural network model includes N encoder layers, and each of the N encoder layers includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism. The first feature tensor includes N value V1 tensors, and obtaining the first feature tensor of the first label sequence based on the first label sequence includes: The tensor obtained based on the first label sequence is processed by the second module in the first encoder layer of the N encoder layers to obtain the first V1 tensor among the N V1 tensors; The output tensor of the second module in the (i-1)th encoder layer of the N encoder layers is processed by the second module in the i-th encoder layer to obtain the i-th V1 tensor among the N V1 tensors, where 1 < i ≤ N.

6. The method according to claim 5, characterized in that, The first association weights include N attention weights, and the inference result obtained based on the first association weights and the first feature tensor for the first data to be inferred includes: The first attention weight and the first V1 tensor among the N attention weights are processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The first attention weight is obtained based on the first query Q tensor and the first key K tensor. The first Q tensor and the first K tensor are obtained by the first module in the first encoder layer processing the tensor obtained based on the first data sequence. The second module in the i-th encoder layer processes the i-th attention weight and the i-th V1 tensor among the N attention weights to obtain the output tensor of the second module in the i-th encoder layer. The i-th attention weight is obtained based on the i-th Q tensor and the i-th K tensor. The i-th Q tensor and the i-th K tensor are obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer. When i equals N, the reasoning result of the first data to be reasoned is obtained based on the output tensor of the second module in the Nth encoder layer.

7. The method according to claim 6, characterized in that, The method further includes: The first module in the first encoder layer processes the first Q tensor, the first K tensor, and the first V2 tensor to obtain the output tensor of the first module in the first encoder layer. The first V2 tensor is obtained by the first module in the first encoder layer processing the tensor based on the first data sequence. The first module in the i-th encoder layer processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor to obtain the output tensor of the first module in the i-th encoder layer. The i-th V2 tensor is obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

8. The method according to claim 6, characterized in that, The first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer. The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor by the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor by the second module in the i-th encoder layer.

9. A method for training a neural network model, characterized in that, The method includes: Obtain a second data sequence and a second label sequence. The second data sequence includes second historical data to be inferred and second data to be inferred. The labels in the second label sequence correspond one-to-one with the data in the second data sequence in terms of sequence position. The following processing is performed through a neural network: a second association weight is obtained between each data in the second data to be inferred and the second historical data to be inferred based on the second data sequence, and an inference result of the second data to be inferred is obtained based on the second association weight and the second label sequence; The parameters of the neural network are adjusted based on the reasoning result of the second data to be reasoned and the label of the second data to be reasoned, so as to obtain the neural network model.

10. The method according to claim 9, characterized in that, The method further includes: The second historical data to be inferred is obtained by sampling the neighborhood of the second data to be inferred in the data feature space.

11. The method according to claim 9, characterized in that, The method further includes: The second target label is obtained by sampling the neighborhood of the label of the second data to be inferred in the label feature space, and the second historical data to be inferred is the historical data to be inferred corresponding to the second target label.

12. The method according to any one of claims 9-11, characterized in that, The reasoning result obtained based on the second association weight and the second label sequence for the second data to be reasoned includes: The second feature tensor of the second label sequence is obtained based on the second label sequence; The reasoning result of the second data to be reasoned is obtained based on the second association weight and the second feature tensor.

13. The method according to claim 12, characterized in that, The neural network includes N encoder layers, and each of the N encoder layers includes a first module based on a self-attention mechanism and a second module based on a self-attention mechanism; The second feature tensor includes N value V1 tensors. The process of obtaining the second feature tensor based on the second label sequence includes: The tensor obtained based on the second label sequence is processed by the second module in the first encoder layer of the N encoder layers to obtain the first V1 tensor among the N V1 tensors; The output tensor of the second module in the (i-1)th encoder layer of the N encoder layers is processed by the second module in the i-th encoder layer to obtain the i-th V1 tensor among the N V1 tensors, where 1 < i ≤ N.

14. The method according to claim 13, characterized in that, The second association weights include N attention weights, and the inference result obtained based on the second association weights and the second feature tensor for the second data to be inferred includes: The first attention weight and the first V1 tensor among the N attention weights are processed by the second module in the first encoder layer to obtain the output tensor of the second module in the first encoder layer. The first attention weight is obtained based on the first query Q tensor and the first key K tensor. The first Q tensor and the first K tensor are obtained by the first module in the first encoder layer processing the tensor obtained based on the second data sequence. The second module in the i-th encoder layer processes the i-th attention weight and the i-th V1 tensor among the N attention weights to obtain the output tensor of the second module in the i-th encoder layer. The i-th attention weight is obtained based on the i-th Q tensor and the i-th K tensor. The i-th Q tensor and the i-th K tensor are obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer. When i equals N, the reasoning result of the second data to be reasoned is obtained based on the output tensor of the second module in the Nth encoder layer.

15. The method according to claim 14, characterized in that, The method further includes: The first module in the first encoder layer processes the first Q tensor, the first K tensor, and the first V2 tensor to obtain the output tensor of the first module in the first encoder layer. The first V2 tensor is obtained by the first module in the first encoder layer processing the tensor based on the second data sequence. The first module in the i-th encoder layer processes the i-th Q tensor, the i-th K tensor, and the i-th V2 tensor to obtain the output tensor of the first module in the i-th encoder layer. The i-th V2 tensor is obtained by the first module in the i-th encoder layer processing the output tensor of the first module in the (i-1)-th encoder layer.

16. The method according to claim 14, characterized in that, The first attention weight is obtained by processing the first Q tensor and the first K tensor through the first module in the first encoder layer, or the first attention weight is obtained by processing the first Q tensor and the first K tensor through the second module in the first encoder layer. The i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor by the first module in the i-th encoder layer, or the i-th attention weight is obtained by processing the i-th Q tensor and the i-th K tensor by the second module in the i-th encoder layer.

17. A data processing apparatus, characterized in that, The device includes a first acquisition unit and a first processing unit, wherein: The first acquisition unit is used to acquire a first data sequence and a first tag sequence. The first data sequence includes first historical data to be inferred and first data to be inferred. The tags in the first tag sequence correspond one-to-one with the data in the first data sequence in terms of sequence position. The first processing unit is configured to obtain a first association weight between each data in the first data to be inferred and the first historical data to be inferred based on the first data sequence; and to obtain the inference result of the first data to be inferred based on the first association weight and the first label sequence.

18. A training device for a neural network model, characterized in that, The device includes a second acquisition unit and a second processing unit, wherein: The second acquisition unit is used to acquire a second data sequence and a second tag sequence. The second data sequence includes second historical data to be inferred and second data to be inferred. The tags in the second tag sequence correspond one-to-one with the data in the second data sequence in terms of sequence position. The second processing unit is configured to perform the following processing through a neural network: obtaining a second association weight between each data in the second data to be inferred and the second historical data to be inferred based on the second data sequence, and obtaining the inference result of the second data to be inferred based on the second association weight and the second label sequence; The second processing unit is further configured to adjust the parameters of the neural network based on the reasoning result of the second data to be reasoned and the label of the second data to be reasoned, so as to obtain the neural network model.

19. An electronic device, characterized in that, The device includes at least one processor coupled to at least one memory for storing one or more computer programs; the at least one processor is configured such that when the electronic device executes the one or more computer programs, it implements the method as claimed in any one of claims 1-8 or 9-16.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for execution by the device, which, when executed, causes the device to perform the method as claimed in any one of claims 1-8 or 9-16.

21. A computer program product, characterized in that, When the computer program product is run by the device, the device performs the method as claimed in any one of claims 1-8 or 9-16.