Data processing method and related equipment

By adding time information to the input of the machine learning model, the problem of model training difficulties caused by the 'garbled sample pair' in the training samples collected by the sliding window is solved, and the accuracy of prediction information is improved.

CN119990356APending Publication Date: 2025-05-13HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311505529.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the training process of machine learning model, there may be a ‘garbage sample pair’ in the training samples collected by the sliding window, making it difficult for the model to learn the mapping relationship between historical data and truth values, thereby reducing the accuracy of prediction information.

Method used

By adding time information to the input of the machine learning model, specifically, the feature information of each information is combined with its corresponding time interval feature information and input it into the machine learning model, thereby distinguishing information from different time periods and improving the model's ability to distinguish different situations.

Benefits of technology

It effectively improves the accuracy of the prediction information output by the machine learning model and solves the problem of model training difficulties caused by the 'garbled sample pair'.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990356A_ABST
    Figure CN119990356A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and related equipment, the method can use an artificial intelligence technology to process sequence data, and the method comprises the following steps: obtaining first information and first time information generated in a first time period, the first time information comprises a relative time interval between a time point for generating the first information and a preset reference time point; the first feature information is input into the machine learning model, prediction information corresponding to a second time period is obtained, the first feature information comprises the first information and feature information of the first time interval, and the second time period is located after the first time period. Even if the training samples A and B carry the same first information and the true values corresponding to the training samples A and B are different, but the first time information corresponding to the training samples A and B is different, the machine learning model can learn the mapping relation between the first information generated in the first time period and the true values. And the accuracy of the obtained prediction information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to a data processing method and related equipment. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0003] Using machine learning models in artificial intelligence technology to process time series data is an application of artificial intelligence. Specifically, historical data of M time units can be input into the machine learning model. The historical data of M time units include multiple information generated within the M time units, so that prediction information corresponding to the future K time units can be generated through the machine learning model. M and K are both positive integers greater than or equal to 1.

[0004] Before training the machine learning model, a large number of training samples need to be collected. When collecting training samples, the window is slid according to the specified sliding time interval to obtain multiple training samples from the historical data. Each training sample may include M+K time units of historical data, where M time units of historical data are used as input to the machine learning model, and K time units of historical data are used to indicate the ground truth corresponding to the aforementioned M time units of historical data. The "ground truth corresponding to the M time units of historical data" can also be understood as the correct label corresponding to the M time units of historical data.

[0005] However, since a sliding window is used to collect training samples, there may be "junk sample pairs" among multiple training samples. For example, the junk sample pair includes training sample A and training sample B. The multiple information included in training sample A is the same as the multiple information included in training sample B, but the true value indicated by training sample A is different from the true value indicated by training sample B. The existence of such junk sample pairs will make it difficult for the machine learning model to learn the mapping relationship between the historical data of M time units and the true value, which in turn causes the accuracy of the prediction information output by the trained machine learning model to be insufficient. Summary of the invention

[0006] The embodiments of the present application provide a data processing method and related equipment to improve the accuracy of prediction information output by a machine learning model.

[0007] The embodiments of the present application provide the following technical solutions:

[0008] In a first aspect, an embodiment of the present application provides a data processing method, which can use artificial intelligence technology to process sequence data, and the method includes: after the execution device obtains at least one first information generated within a first time period, it can also obtain first time information, and the first time information includes a first time interval between the time point when the first information is generated and the first time point; it should be understood that the "first time point" can be understood as a preset reference time point, and the "first time interval between the time point when the first information is generated and the first time point" can also be understood as the relative time interval between the time point when the first information is generated and the first time point (that is, the preset reference time point).

[0009] The execution device can obtain the first feature information based on the at least one first information and the first time information, and then input the first feature information into the machine learning model to obtain the prediction information corresponding to the second time period output by the machine learning model. Among them, the first feature information includes the feature information of each first information and the feature information of the first time interval corresponding to each first information. Further, "the feature information of the first information" can be understood as the information obtained after performing a feature extraction operation on the first information, and "the feature information of the first time interval" can be understood as the information obtained after performing a feature extraction operation on the first time interval. However, it should be noted that the aforementioned feature extraction operation can be performed by other devices other than the execution device, or it can be performed by the execution device, which can be determined in combination with actual conditions.

[0010] Exemplarily, the first feature information may include initial feature information of each first information and initial feature information of the first time interval corresponding to each first information; after the execution device inputs the first feature information into the machine learning model, it updates and processes the first feature information through the machine learning model, and obtains the prediction information corresponding to the second time period output by the machine learning model.

[0011] Exemplarily, the second time period may overlap with the first time period, that is, there is an overlapping time period between the first time period and the second time period; or, there may be no overlap between the second time period and the first time period, that is, there is no overlapping time period between the first time period and the second time period. When there is no overlap between the second time period and the first time period, in one case, the second time period and the first time period may be adjacent, that is, the end time point of the first time period may be the start time point of the second time period; or, in another case, there is an interval between the first time period and the second time period, that is, after the end time point of the first time period, it may be a period of time before the start time point of the second time period; it should be understood that the "relationship between the first time period and the second time period" can be flexibly set in combination with the actual application scenario.

[0012] In this implementation, the input of the machine learning model includes not only each first information generated within the first time period, but also first time information, where the first time information indicates the relative time interval between the time point at which each first information is generated and a preset reference time point. When the first information generated in two different first time periods (for example, first time period 1 and first time period 2) is the same, the first time information corresponding to the first information generated in first time period 1 may be different from the first time information corresponding to the first information generated in first time period 2. That is, even if the first information generated in two different first time periods is the same, the two different first time periods can be distinguished with the help of the aforementioned preset reference time point, thereby facilitating the machine learning model to better distinguish different situations, which is conducive to improving the accuracy of the prediction information output by the machine learning model.

[0013] In a possible implementation, the first time point may be any time point with a fixed relative time interval from the first time period, for example, the first time point may be determined based on the start time point or the end time point of the first time period; illustratively, the first time point may be within the first time period or outside the first time period. Optionally, the first time point may specifically adopt the start time point or the end time point of the first time end. Or, for example, the first time point may be a time point located in the middle of the first time period, and the aforementioned time point located in the middle is determined based on the start time point and the end time point of the third time period. For another example, the first time point is 5 minutes earlier than the start time point of the first time period. For another example, the first time point is 6 minutes earlier than the end time point of the first time period. For another example, the first time point is 5 minutes later than the end time point of the first time period, etc.

[0014] In this implementation, the first time point is determined based on the start time point or the end time point of the third time period, which provides a scheme for determining the first time point and reduces the difficulty of implementing this scheme; optionally, when the first time point adopts the start time point or the end time point of the third time period, the difficulty of the process of "determining the first time point" is further reduced.

[0015] In a possible implementation, the first information includes alarm information of the network device, the aforementioned alarm information may include an alarm type, the prediction information corresponding to the second time period indicates whether the network device has a first fault in the second time period, and the time point of generating the first information is the time point of occurrence of the alarm information. For example, the aforementioned network device may be a base station, a network device in a core network, or a network device in other forms, etc.; for example, "alarm type" can also be understood as "alarm name", and the aforementioned alarm type includes but is not limited to the radio frequency unit Zhu Bo alarm, optical network unit (optical network unit, ONU) signal loss, or low temperature, etc., and the examples here are only for the convenience of understanding the concept of alarm type; for example, the first fault may include any one or more of the following faults: base station decommissioning fault, cell decommissioning fault, optical power fault, single board temperature fault, combined wave link interruption fault or optical path interruption fault, etc., and the specific manifestation of the "first fault" can be determined in combination with the actual application scenario. In this implementation, a specific application scenario of the present application is provided, which improves the degree of integration between the present solution and the specific application scenario.

[0016] In one possible implementation, the first information includes the user's physiological information, the prediction information corresponding to the second time period indicates the user's motion state in the second time period, and the time point when the first information is generated is the time point when the physiological information occurs. Exemplarily, the user's physiological information includes but is not limited to: heart rate, arm swing amplitude or other physiological information, etc.; the user's motion state includes but is not limited to: walking state, running state, going upstairs state, going downstairs state or other motion state, etc. In this implementation, another specific application scenario of the present application is provided, which improves the implementation flexibility of the present solution.

[0017] In one possible implementation, the first feature information includes at least one second feature information corresponding to at least one first information, and each second feature information includes feature information of a first information and feature information of a first time interval corresponding to a first information. In this method, in the process of the execution device updating the feature of the first feature information through the machine learning model, at least one second feature information is divided into at least two groups based on each of at least two grouping criteria; the execution device can update the second feature information in the same group according to the second feature information in the same group of the aforementioned at least two groups, that is, when the execution device updates the feature of a second feature information in any one of the aforementioned at least two groups (hereinafter referred to as the "target group" for convenience of description), only the second feature information in the target group is used, and the second feature information of other groups outside the target group will not be used.

[0018] Optionally, the execution device may update the second feature information in the same group based on the attention mechanism according to the second feature information in the same group of at least two groups.

[0019] In this implementation, multiple second feature information are grouped based on at least two different perspectives, and then the second feature information in the same group after grouping is updated, so that the feature update process of the first feature information can be very sufficient, which is conducive to the machine learning model learning rich information, thereby helping to improve the accuracy of the prediction information output by the machine learning model.

[0020] In one possible implementation, if the first information is alarm information including an alarm type, and the prediction information corresponding to the second time period indicates whether a first fault occurs within the second time period, then the first characteristic information (or the updated first characteristic information) may include second characteristic information corresponding one-to-one to at least one alarm information, each second characteristic information including characteristic information of an alarm information and characteristic information of a first time interval corresponding to the aforementioned alarm information. Since each second characteristic information corresponds one-to-one to an alarm information, at least two grouping criteria include any at least two of the following: the generation time of the first information (i.e., the alarm information), whether the alarm type included in the first information (i.e., the alarm information) can directly indicate the first fault, or the alarm type in the first information (i.e., the alarm information).

[0021] Exemplarily, "the first alarm type can directly indicate the first fault" can be understood as when it is determined that the first alarm type occurs, it means that the first fault will occur; for example, when the first fault is a base station decommissioning fault, the first alarm type may include a base station out of service (gNodeB Out of Service), a communication service company error (CSL Fault) or other alarm types, etc.; for another example, when the first fault is a rectifier fault, the first alarm type may include a rectifier module fault (rectifier module fault), a rectifier fault (Rectifier Fault) or other types of alarm types, etc.

[0022] In this implementation, a variety of specific grouping criteria are provided to improve the implementation flexibility of this solution; in addition, the grouping criteria include the time when the alarm information is generated, whether the alarm information can directly indicate the first fault, and the type of alarm carried by the alarm information. These grouping criteria are very suitable for the application scenario of "the first information is the alarm information", which is more conducive to the machine learning model to obtain sufficient information, and is conducive to improving the accuracy of the prediction information output by the machine learning model.

[0023] In one possible implementation, the first information is warning information, and the prediction information corresponding to the second time period indicates whether the first fault occurs in the second time period. The execution device updates the second characteristic information in the same group according to the second characteristic information in the same group of at least two groups, including any at least two of the following multiple implementations:

[0024] In one implementation, when the first grouping basis among at least two grouping bases is adopted, the at least two groups can be at least two first groups; wherein, based on the first grouping basis, the first time period can be divided into at least two sub-time periods, and each first group among the at least two first groups corresponds to a sub-time period among the at least two sub-time periods, that is, each first group includes second characteristic information corresponding to all first information generated within a sub-time period among the at least two sub-time periods. Then, the execution device can update each second characteristic information within the same first group according to the second characteristic information within the same first group among the at least two first groups, wherein, based on the first grouping basis. Or,

[0025] In another implementation, when the second grouping basis among the at least two grouping bases is adopted, the at least two groups include a second group and a third group, wherein the second group includes second characteristic information corresponding to all first information in at least one first information that can directly indicate a first fault, and the third group includes second characteristic information corresponding to all first information other than the second group in at least one first information. The execution device updates each second characteristic information in the second group according to the second characteristic information in the second group, and updates each second characteristic information in the third group according to the second characteristic information in the third group. Or,

[0026] In another implementation, when the third grouping basis among the at least two grouping basis is adopted, the at least two groups may be at least two fourth groups, wherein the alarm type of the first information corresponding to all the second characteristic information in each fourth group is the same. Then the execution device may update each second characteristic information in the same fourth group according to the second characteristic information in the same fourth group among the at least two fourth groups.

[0027] In this implementation method, specific implementation steps for updating the features of the first feature information using a machine learning model are provided, which is conducive to reducing the difficulty of implementing this solution.

[0028] In a possible implementation, when the first information is specifically expressed as alarm information, the execution device may pre-store second information, wherein the second information indicates feature information corresponding to each of the multiple alarm information, and the second information further indicates feature information corresponding to each of the multiple time intervals, and the second information is updated during the execution of the training operation of the machine learning model. Exemplarily, the feature information corresponding to the aforementioned alarm information is obtained based on a feature extraction operation performed on the alarm information, and the feature information corresponding to the aforementioned time interval is obtained based on a feature extraction operation performed on the time interval.

[0029] The execution device may also obtain the feature information of each first information and the feature information of the first time interval corresponding to each first information based on the pre-stored second information before inputting the first feature information into the machine learning model to obtain the prediction information corresponding to the second time period output by the machine learning model, so as to obtain the first feature information. It should be noted that in this implementation, since the feature information of the first information and the feature information of the first time interval are directly obtained by the execution device from the second information, the above-mentioned feature extraction operation is not performed by the execution device.

[0030] In this implementation, second information is deployed in the first device, and the second information includes characteristic information of each alarm type and characteristic information of each time interval. The initial characteristic information corresponding to the first sample (that is, the first characteristic information) can be quickly acquired based on the second information, and the second information will also be updated during the training phase of the machine learning model, which is conducive to obtaining more accurate initial characteristic information of the first sample.

[0031] In one possible implementation, the training process of the machine learning model uses multiple training samples, a true value corresponding to each training sample, and a loss function, each training sample includes at least one first information generated within a third time period, and the true value corresponding to the training sample indicates whether a first fault occurs within a fourth time period, and the fourth time period is located after the third time period.

[0032] Among them, when the true value corresponding to the training sample indicates that the first fault occurred within the fourth time period, the training sample is a positive training sample, and the above-mentioned loss function is determined based on the characteristic information of the positive training sample. The goal of training using the loss function includes improving the similarity between the characteristic information of different positive training samples.

[0033] In this implementation, the second loss function is used to improve the similarity between the feature information of different positive training samples generated by the machine learning model during the feature update process, thereby improving the machine learning model's ability to understand different positive samples, which is conducive to further improving the accuracy of the prediction information generated by the machine learning model.

[0034] In a possible implementation, when the true value corresponding to the training sample indicates that the first fault did not occur within the fourth time period, the training sample is a negative training sample, and the aforementioned loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample. The goal of training using the loss function also includes: improving the similarity between the feature information of different negative training samples, and / or reducing the similarity between the feature information of the positive training sample and the feature information of the negative training sample.

[0035] In this implementation, using the second loss function to train the machine learning model can also improve the similarity between the feature information of different negative training samples, and reduce the similarity between the feature information of positive training samples and the feature information of negative training samples, which is beneficial to improving the machine learning model's ability to understand different negative samples, and is beneficial to improving the machine learning model's ability to distinguish between positive samples and negative samples, and is beneficial to further improving the accuracy of the prediction information generated by the machine learning model.

[0036] In a second aspect, an embodiment of the present application provides a data processing method, which can use artificial intelligence technology to process sequence data, and the method includes: a training device obtains a training sample, the training sample includes at least one first information generated within a third time period; obtains first time information, the first time information includes a first time interval between a time point when the first information is generated and the first time point, the first time point is a preset reference time point, and the first time interval is a relative time interval between a time point when the first information is generated and the first time point; inputs the first feature information into a machine learning model to obtain prediction information corresponding to a fourth time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of the first time interval corresponding to the first information, and the fourth time period is located after the third time period; trains the machine learning model according to a first loss function, and the first loss function indicates the similarity between the prediction information corresponding to the fourth time period and the true value corresponding to the training sample.

[0037] In this implementation, since the input of the machine learning model includes not only each first information generated within the first time period, but also includes the first time information, and the first time information indicates the relative time interval between the time point at which each first information is generated and the preset reference time point, then in the training stage of the machine learning model, when the machine learning model encounters a "junk sample pair", although the multiple first information generated in the historical data of M time units included in training sample A and training sample B are the same, since the training samples are collected using a sliding window, the relative time interval between the time point at which the multiple first information is generated in training sample A and the first time point is different from the relative time interval between the multiple first information generated in training sample B and the first time point. After the relative time interval is added to the input of the machine learning model, training sample A and training sample B are different for the machine learning model. Then, when the true values ​​included in training sample A and training sample B are different, the machine learning model can also learn the mapping relationship between the historical data and the true value within the first time period.

[0038] In the second aspect of the present application, the training device can also execute the steps performed by the execution device in each possible implementation method of the first aspect. The meanings of the nouns in the second aspect and each possible implementation method of the second aspect and the beneficial effects brought about can all be referred to the first aspect and will not be repeated here.

[0039] In a third aspect, an embodiment of the present application provides a data processing method, which can use artificial intelligence technology to process sequence data, and the method includes: executing a device to obtain at least one alarm information generated within a first time period; inputting the first feature information into a machine learning model, the first feature information including at least one second feature information corresponding to the at least one alarm information; updating the second feature information in the same group of at least two groups according to the second feature information in the same group, and obtaining prediction information corresponding to the second time period output by the machine learning model, wherein, in the process of updating the second feature information through the machine learning model, based on each of the at least two grouping bases, the at least one second feature information is divided into at least two groups, and the second time period is located after the first time period.

[0040] In one possible implementation, the method also includes: executing a device to obtain first time information, the first time information includes a first time interval between a time point when the alarm information is generated and the first time point, the first time point is a preset reference time point, the first time interval is a relative time interval between the time point when the alarm information is generated and the first time point, and each second characteristic information includes characteristic information of an alarm information and characteristic information of a first time interval corresponding to an alarm information.

[0041] In the third aspect of the present application, the execution device can also execute the steps executed by the execution device in each possible implementation method of the first aspect. The meanings of the terms in the third aspect and each possible implementation method of the third aspect and the beneficial effects brought about can all be referred to the first aspect and will not be repeated here.

[0042] In a fourth aspect, an embodiment of the present application provides a data processing device, which can use artificial intelligence technology to process sequence data, and the data processing device includes: an acquisition module, used to acquire at least one first information generated within a first time period; the acquisition module is also used to acquire first time information, the first time information includes a first time interval between a time point when the first information is generated and the first time point, and the first time point is a preset reference time point; an input module, used to input the first feature information into a machine learning model to obtain prediction information corresponding to a second time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of a first time interval corresponding to the first information, and the second time period is located after the first time period.

[0043] In a possible implementation, the first time point is determined based on a start time point or an end time point of the first time period.

[0044] In one possible implementation, the first information includes alarm information of the network device, the prediction information corresponding to the second time period indicates whether a first fault occurs in the network device within the second time period, and the time point when the first information is generated is the time point when the alarm information occurs.

[0045] In one possible implementation, the first information includes physiological information of the user, the prediction information corresponding to the second time period indicates the motion state of the user in the second time period, and the time point when the first information is generated is the time point when the physiological information occurs.

[0046] In one possible implementation, the first feature information includes at least one second feature information corresponding to at least one first information, and each second feature information includes feature information of one first information and feature information of a first time interval corresponding to one first information. In the process of updating the features of the first feature information through the machine learning model, based on each of at least two grouping criteria, at least one second feature information is divided into at least two groups, and the device further includes: an updating module, which is used to update the second feature information in the same group according to the second feature information in the same group of the at least two groups.

[0047] In one possible implementation, the first information includes an alarm type, and the prediction information corresponding to the second time period indicates whether a first fault occurs within the second time period; the at least two grouping criteria include any at least two of the following: the generation time of the first information, whether the alarm type included in the first information can directly indicate the first fault, or the alarm type in the first information.

[0048] In one possible implementation, the first information is warning information, and the prediction information corresponding to the second time period indicates whether the first fault occurs in the second time period. The step of updating the second characteristic information in the same group according to the second characteristic information in the same group of at least two groups includes any of the following:

[0049] updating the second characteristic information in the same first group according to the second characteristic information in the same first group in the at least two first groups, wherein when a first grouping basis in the at least two grouping basis is adopted, the at least two groups are at least two first groups, based on the first grouping basis, the first time period is divided into at least two sub-time periods, and each first group includes the second characteristic information corresponding to all the first information generated in one sub-time period in the at least two sub-time periods; or,

[0050] The second characteristic information in the second group is updated according to the second characteristic information in the second group, and the second characteristic information in the third group is updated according to the second characteristic information in the third group, wherein when the second grouping basis among the at least two grouping basis is adopted, the at least two groups include the second group and the third group, the second group includes the second characteristic information corresponding to all the first information in the at least one first information that can directly indicate the first fault, and the third group includes the second characteristic information corresponding to all the first information in the at least one first information except the second group; or,

[0051] According to the second characteristic information in the same fourth group in at least two fourth groups, the second characteristic information in the same fourth group is updated, wherein, when the third grouping basis among at least two grouping basis is adopted, at least two groups are at least two fourth groups, and the alarm type of the first information corresponding to all the second characteristic information in each fourth group is the same.

[0052] In one possible implementation, the acquisition module is also used to acquire characteristic information of the first information and characteristic information of the first time interval corresponding to the first information based on pre-stored second information to obtain the first characteristic information, wherein the second information indicates characteristic information corresponding to each of multiple alarm information, and the second information also indicates characteristic information corresponding to each of multiple time intervals, and the second information is updated during the process of executing the training operation of the machine learning model; the characteristic information corresponding to the alarm information is obtained based on the feature extraction operation performed on the alarm information, and the characteristic information corresponding to the time interval is obtained based on the feature extraction operation performed on the time interval.

[0053] In one possible implementation, the training process of the machine learning model uses multiple training samples, a true value corresponding to each training sample, and a loss function, each training sample includes at least one first information generated within a third time period, and the true value corresponding to the training sample indicates whether a first fault occurred within a fourth time period, and the fourth time period is located after the third time period; wherein, when the true value corresponding to the training sample indicates that a first fault occurred within the fourth time period, the training sample is a positive training sample, the loss function is determined based on the feature information of the positive training sample, and the goal of training using the loss function includes improving the similarity between the feature information of different positive training samples.

[0054] In one possible implementation, when the true value corresponding to the training sample indicates that the first fault did not occur within a fourth time period, the training sample is a negative training sample, and the loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample. The goal of training using the loss function also includes: improving the similarity between the feature information of different negative training samples, and / or reducing the similarity between the feature information of the positive training sample and the feature information of the negative training sample.

[0055] The specific implementation methods, meanings of terms and beneficial effects of the steps in each possible implementation method of the fourth aspect of the present application can all be referred to the first aspect and will not be repeated here.

[0056] In a fifth aspect, an embodiment of the present application provides a data processing device, which can use artificial intelligence technology to process sequence data, and the data processing device includes: an acquisition module, used to acquire training samples, the training samples include at least one first information generated within a third time period; the acquisition module is also used to acquire first time information, the first time information includes a first time interval between a time point when the first information is generated and the first time point, and the first time point is a preset reference time point; an input module, used to input the first feature information into a machine learning model to obtain prediction information corresponding to a fourth time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of the first time interval corresponding to the first information, and the fourth time period is located after the third time period; a training module, used to train the machine learning model according to a first loss function, the first loss function indicates the similarity between the prediction information corresponding to the fourth time period and the true value corresponding to the training sample.

[0057] In one possible implementation, the first feature information includes at least one second feature information corresponding one-to-one to at least one first information, and each second feature information includes feature information of a first information and feature information of a first time interval corresponding to the first information; wherein, in the process of updating the features of the first feature information through a machine learning model, based on each of at least two grouping criteria, at least one second feature information is divided into at least two groups, and the data processing device also includes: an update module, which is used to update the feature information of the first information in the same group according to the feature information of the first information in the same group of at least two groups.

[0058] In one possible implementation, the first information is alarm information, and the prediction information corresponding to the fourth time period indicates whether a first fault occurs within the fourth time period; the acquisition module is also used to obtain characteristic information of the first information and characteristic information of a first time interval corresponding to the first information based on the second information, wherein the second information indicates characteristic information corresponding to each type of alarm information among multiple alarm information, and the second information also indicates characteristic information corresponding to each time interval among multiple time intervals, the characteristic information corresponding to the alarm information is obtained based on a feature extraction operation performed on the alarm information, and the characteristic information corresponding to the time interval is obtained based on a feature extraction operation performed on the time interval; the training module is specifically used to update the weight parameters and the second information of the machine learning model according to the first loss function.

[0059] In one possible implementation, when the true value corresponding to the training sample indicates that a first fault occurred within a fourth time period, the training sample is a positive training sample, wherein the training module is specifically used to train the machine learning model according to a first loss function and a second loss function, the second loss function is determined based on feature information of the positive training sample, and the goal of training using the second loss function includes improving the similarity between feature information of different positive training samples.

[0060] In one possible implementation, when the true value corresponding to the training sample indicates that the first fault did not occur within a fourth time period, the training sample is a negative training sample, wherein the second loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample, and the goal of training using the second loss function also includes: improving the similarity between the feature information of different negative training samples, and / or reducing the similarity between the feature information of the positive training sample and the feature information of the negative training sample.

[0061] The specific implementation methods of the steps in each possible implementation method of the fifth aspect of the present application, the meaning of the terms and the beneficial effects brought about can all be referred to the second aspect and will not be repeated here.

[0062] In a sixth aspect, an embodiment of the present application provides a data processing device, which can use artificial intelligence technology to process sequence data, and the data processing device includes: an acquisition module, used to acquire at least one alarm information generated within a first time period; an input module, used to input first feature information into a machine learning model, the first feature information including at least one second feature information corresponding to the at least one alarm information; an update module, used to update the second feature information in the same group according to the second feature information in the same group of at least two groups, and obtain prediction information corresponding to the second time period output by the machine learning model, wherein, in the process of updating the second feature information through the machine learning model, based on each of the at least two grouping criteria, at least one second feature information is divided into at least two groups, and the second time period is located after the first time period.

[0063] In one possible implementation, the acquisition module is also used to obtain first time information, the first time information includes a first time interval between a time point when the alarm information is generated and the first time point, the first time point is a preset reference time point, and each second characteristic information includes characteristic information of an alarm information and characteristic information of a first time interval corresponding to an alarm information.

[0064] The specific implementation methods of the steps in each possible implementation method of the sixth aspect of the present application, the meaning of the terms and the beneficial effects brought about can all be referred to the third aspect and will not be repeated here.

[0065] In the seventh aspect, an embodiment of the present application provides a device, including a processor and a memory, where the processor is coupled to the memory, the memory is used to store programs; the processor is used to execute the programs in the memory, so that the device executes the data processing method of the first aspect, the second aspect or the third aspect mentioned above.

[0066] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the data processing method of the first aspect, the second aspect or the third aspect mentioned above.

[0067] In a ninth aspect, an embodiment of the present application provides a computer program product, which includes a program. When the program runs on a computer, the computer executes the data processing method of the first aspect, the second aspect or the third aspect mentioned above.

[0068] In a tenth aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for a terminal device or a communication device. The chip system can be composed of a chip, or it can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1a A schematic diagram of an architecture of a wireless communication system provided in an embodiment of the present application;

[0070] Figure 1b Another schematic diagram of the architecture of a wireless communication system provided in an embodiment of the present application;

[0071] Figure 2 A schematic diagram of a garbage sample pair provided in an embodiment of the present application;

[0072] Figure 3 A flowchart of a data processing method provided in an embodiment of the present application;

[0073] Figure 4 Another schematic diagram of a data processing method provided in an embodiment of the present application;

[0074] Figure 5 A schematic diagram of the first time information provided by an embodiment of the present application;

[0075] Figure 6 A schematic diagram of grouping at least one second characteristic information provided in an embodiment of the present application;

[0076] Figure 7 Another schematic diagram of grouping at least one second feature information provided in an embodiment of the present application;

[0077] Figure 8 Another schematic diagram of grouping at least one second feature information provided in an embodiment of the present application;

[0078] Fig. 9 A schematic diagram of updating features through a first machine learning model provided in an embodiment of the present application;

[0079] Fig.10 A schematic diagram of multiple information indicated by the second loss function provided in an embodiment of the present application;

[0080] Fig.11 Another flowchart of the data processing method provided in the embodiment of the present application;

[0081] Fig.12 A schematic diagram of a second sample provided in an embodiment of the present application;

[0082] Fig.13 Another schematic diagram of a data processing method provided in an embodiment of the present application;

[0083] Fig.14 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;

[0084] Fig.15 Another structural schematic diagram of a data processing device provided in an embodiment of the present application;

[0085] Fig.16 Another structural schematic diagram of a data processing device provided in an embodiment of the present application;

[0086] Fig.17 A schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0087] Fig.18 It is a structural schematic diagram of a training device provided in an embodiment of the present application;

[0088] Fig.19 A schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0089] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0090] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0091] "Send" and "receive" in the embodiments of the present application indicate the direction of signal transmission. For example, "send information to XX device" can be understood as the destination of the information is XX device, which can include direct transmission through the air interface, and also include indirect transmission through the air interface by other units or modules. "Receive information from YY device" can be understood as the source of the information is YY device, which can include direct reception from YY device through the air interface, and can also include indirect reception from YY device through other units or modules through the air interface. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be carried out between devices or within the device, for example, sending or receiving between components, modules, chips, software modules or hardware modules within the device through a bus, a wiring or an interface. It is understandable that the information may be processed as necessary between the source and the destination of the information transmission, such as encoding, modulation, etc., but the destination can understand the valid information from the source. Similar expressions in this application can be understood in a similar way and will not be repeated.

[0092] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information (such as the indication information described below) is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, directly indicating the information to be indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated may also be indirectly indicated by indicating other information, wherein there is an association between the other information and the information to be indicated; it may also be possible to indicate only a part of the information to be indicated, while the other part of the information to be indicated is known or agreed in advance, for example, the indication of specific information can be realized by means of the arrangement order of each information agreed in advance (such as predefined by the protocol), thereby reducing the indication overhead to a certain extent. The present application does not limit the specific method of indication. It is understandable that, for the sender of the indication information, the indication information can be used to indicate the information to be indicated, and for the receiver of the indication information, the indication information can be used to determine the information to be indicated.

[0093] The present application may use artificial intelligence technology to process sequential data, where sequential data refers to data with a sequence. Optionally, the sequence of the aforementioned serialized data may be determined based on the time when the data was generated.

[0094] Exemplarily, in an application scenario, the method provided in the present application can be applied to the field of communications, and the above-mentioned serialized data may include multiple alarm information generated within a period of time. At least one alarm information generated within the first time period is named as a first sample. Based on the aforementioned first sample, it is processed by a machine learning model (for the convenience of description, it is hereinafter referred to as the "first machine learning model") to obtain prediction information corresponding to the second time period output by the first machine learning model, and the prediction information corresponding to the second time period indicates whether a first fault has occurred in the second time period.

[0095] To more intuitively understand the application scenarios of the method provided in this application, please refer to Figure 1a , Figure 1a The method provided in this application can be applied to a wireless communication system, such as Figure 1a As shown, the wireless communication system includes a network device 101 and a terminal device (mobile station, MS) 102. A wireless connection can be established between the network device 101 and each terminal device 102, and a wireless connection can also be established between each terminal device 102.

[0096] The network device 101 may refer to a device that provides wireless access services in a wireless network. Exemplarily, the network device 101 may be a device that connects the terminal device 102 to the wireless network, and may also be called a base station; the aforementioned base station may be various forms of macro base stations, micro base stations, relay stations or access points, etc. In wireless communication systems using different wireless access technologies, the name of the network device 101 having the base station function may be different. For example, the base station may be called an evolved Node B (eNB), a Node B (NB), the next generation Node B (gNB) in the fifth generation (5G) communication system, a home base station (e.g., home evolved NodeB, or home Node B, HNB), a base band unit (BBU), a wireless fidelity (Wi-Fi) access point (AP), a transmission reception point (TRP) or a radio network controller (RNC), etc. In another possible scenario, multiple network nodes collaborate to assist in achieving wireless access, and different network nodes respectively implement part of the functions of the base station. For example, a network node may be a centralized unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU), etc. CU and DU may be separately configured, or may be included in the same network element, such as a baseband unit (BBU). RU may be included in a radio frequency device or a radio frequency unit, such as a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH). In different systems, CU (or CU-CP and CU-UP), DU, or RU may also have different names, but those skilled in the art may understand their meanings. For example, in the ORAN system, CU may also be called open CU (O-CU), DU may also be called open DU (O-DU), CU-CP may also be called open CU-CP (O-CU-CP), CU-UP may also be called open CU-UP (O-CU-UP), and RU may also be called open RU (O-RU).Among them, any unit in CU (or CU-CP, CU-UP), DU and RU can be implemented by software module, hardware module, or a combination of software module and hardware module. The embodiment of the present application does not limit the specific device form of the network device 101.

[0097] The terminal device 102 refers to a wireless terminal device that can receive scheduling information and indication information sent by the network device 101. The terminal device 102 can be a handheld device with wireless communication function, a vehicle-mounted device, a wearable device, a computing device or other processing device, etc., which are not exhaustive here.

[0098] The terminal device 102 can communicate with one or more core networks or the Internet via a wireless access network (RAN). For example, the terminal device 102 can be a portable, pocket-sized, handheld, computer-built-in or vehicle-mounted mobile device that exchanges voice and / or data with the wireless access network. Exemplarily, the terminal device 102 can be a user agent, a cellular phone, a smart phone, a personal digital assistant (PDA), a tablet PC (Tablet Personal Computer, Tablet PC), a wireless modem, a handset, a laptop computer, a personal communication service (PCS) phone, a remote station (remotestation), an access point (AP), a remote terminal device (remote terminal), an access terminal device (access terminal), a customer premises equipment (CPE), a terminal, a user equipment (UE) or a mobile terminal (MT), etc.

[0099] For another example, the terminal device 102 may also be a wearable device, which is a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not just hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various types of smart bracelets, smart helmets, and smart jewelry for vital sign monitoring.

[0100] For another example, the terminal device 102 may also be a drone, a robot, a terminal device in device-to-device (D2D) communication, a terminal device in vehicle to everything (V2X), a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a terminal device in self driving, a terminal device in remote medical, a terminal device in a smart grid, a wireless terminal in a smart city, a terminal device in a smart home, etc.

[0101] In addition, the terminal device 102 may also be a terminal device in a communication system after the 5G communication system (for example, a sixth generation (6G) communication system, etc.) or a terminal device in a future evolved public land mobile network (PLMN), etc. The embodiment of the present application does not limit the device form of the terminal device 102.

[0102] Exemplarily, the network device 101 can generate alarm information during operation, and at least one alarm information generated by the network device 101 within a first time period can be obtained (that is, a first sample is obtained). Based on the first sample, the first machine learning model is used for processing to obtain prediction information corresponding to the second time period output by the first machine learning model. The prediction information corresponding to the second time period indicates whether the network device 101 will have a certain fault within the second time period.

[0103] Please continue reading Figure 1b , Figure 1bAnother schematic diagram of the architecture of the wireless communication system provided in the embodiment of the present application. Figure 1b As shown in FIG. 1 , in a smart home scenario, various smart home products are connected via a wireless network to enable data to be transmitted between the smart home products. Figure 1b In the embodiment, smart home products such as smart TV, smart air purifier, smart water dispenser, smart speaker and sweeping robot are taken as examples. These smart home products are connected to the same wireless network through wireless router 201, so as to realize data interaction between various smart home products. In addition to the smart home products in the above examples, other types of smart home products may also be included in actual applications, such as smart refrigerators, smart range hoods, smart curtains and other smart home products. This embodiment does not limit the types of smart home products.

[0104] Exemplarily, the wireless router 201 can generate alarm information during operation, and at least one alarm information generated by the wireless router 201 within a first time period can be obtained, and prediction information corresponding to the second time period can be generated through the first machine learning model. The prediction information corresponding to the second time period indicates whether the wireless router 201 will have a certain fault within the second time period.

[0105] In addition to the above Figure 1a and Figure 1b In addition to the scenarios described above, the method provided in the embodiments of the present application can also be applied to other communication system scenarios. For example, the network elements in the core network and the base stations can also be connected through a wireless network and transmit data to each other through the wireless network. Then, based on the alarm information generated by the network elements of the core network in the first time period, the first machine learning model can be used to predict whether the network elements of the core network will have a certain fault in the second time period. The embodiments of the present application do not limit the specific scenarios to which the method provided in the present application is applied.

[0106] It should be noted that the wireless communication systems mentioned in the embodiments of the present application include but are not limited to: fifth generation mobile communication technology (5th Generation Mobile Communication Technology, 5G) communication system, 6G communication system, satellite communication system, short-range communication system, narrowband Internet of Things system (Narrow Band-Internet of Things, NB-IoT), Global System for Mobile Communications (Global System for Mobile Communications, GSM), Enhanced Data rate for GSM Evolution (Enhanced Data rate for GSM Evolution, EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access 2000 (Code Division Multiple Access, CDMA2000), Time Division-Synchronization Code Division Multiple Access (TD-SCDMA) and Long Term Evolution (LTE) and other communication systems. The embodiments of the present application do not limit the specific architecture of the wireless communication system.

[0107] For another example, in another application scenario, the method provided by this application can be applied to the field of smart terminals, and the above serialized data may include physiological information generated by the user within a period of time. The physiological information generated by the user within a first time period is collected by a smart wearable device, and prediction information corresponding to the second time period is generated based on the first machine learning model. The prediction information corresponding to the second time period indicates the user's motion state within the second time period, etc. It should be noted that the above introduction to the application scenario of the method provided by this application is only for the convenience of understanding this solution. The method provided by this application can also be applied to other scenarios, which are not listed one by one in the embodiments of this application.

[0108] The first machine learning model will be used in the above-mentioned application scenarios. The "training phase of the first machine learning model" is the process of training the first machine learning model once or multiple times using training samples; the "inference phase of the first machine learning model" is the process of using the first machine learning model to process data. The process of iteratively training the first machine learning model is also the process of iteratively updating the parameters in the first machine learning model. After iteratively training the first machine learning model using training samples, the trained parameters corresponding to the first machine learning model can be obtained. The aforementioned parameters obtained in the training phase will be used in the inference phase.

[0109] However, there may be junk sample pairs in the training data used to train the first machine learning model. For a more intuitive understanding of the concept of "junk sample pairs", please refer to Figure 2 , Figure 2 A schematic diagram of a junk sample pair provided in an embodiment of the present application, wherein the junk sample pair includes a training sample A and a training sample B, such as Figure 2 As shown, the multiple information generated in the historical data of M time units included in the training sample A is the same as the multiple information generated in the historical data of M time units included in the training sample B. When the true value (ground truth) indicated when the B information appears within K time units is different from the true value indicated when the B information does not appear within K time units, the true value in the training sample A is different from the true value in the training sample B; it should be understood that the "true value" can also be replaced by "correct label", "expected information" or other names, etc., which are not limited in the embodiments of the present application. Figure 2 The parts of the training samples A and B shown in the figure that are input into the machine learning model (that is, the multiple information generated within M time units) are the same, but correspond to different true values. The existence of such junk sample pairs will make it difficult for the machine learning model to learn the mapping relationship between the historical data of M time units and the true value, which in turn leads to the accuracy of the prediction information output by the machine learning model is not high enough.

[0110] In combination with the above-described problems, the present application provides a data processing method. The method in the present application not only provides the steps of the training phase of the first machine learning model, but also provides the steps of the reasoning phase of the first machine learning model. Figure 3 The steps of the inference phase of the first machine learning model are described as follows Figure 3 The steps in the corresponding embodiments can be executed by an execution device. For details, please refer to Figure 3 , Figure 3 A flow chart of a data processing method provided in an embodiment of the present application.

[0111] 301. Obtain at least one first information generated within a first time period.

[0112] Exemplarily, the first information may be alarm information of a network device, or the first information may be physiological information of a user, or the first information may also carry other types of information, etc., which are not exhaustively listed here.

[0113] 302. Acquire first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point.

[0114] Exemplarily, the first time point may be any time point with a fixed relative time interval from the first time period; the first time point may be within the first time period or outside the first time period. For example, the first time point may be the starting time point of the first time period. For another example, the first time point may be the ending time point of the first time period. For another example, the first time point may be a time point located in the middle of the first time period, and the aforementioned time point located in the middle is determined based on the starting time point and the ending time point of the third time period. For another example, the first time point is 5 minutes earlier than the starting time point of the first time period. For another example, the first time point is 6 minutes earlier than the ending time point of the first time period. For another example, the first time point is 5 minutes later than the ending time point of the first time period, and so on, which are not exhaustive here.

[0115] 303. Input the first feature information into the first machine learning model to obtain prediction information corresponding to the second time period output by the first machine learning model, wherein the first feature information includes feature information of the first information and feature information of the first time interval corresponding to the first information, and the second time period is located after the first time period.

[0116] Exemplarily, the first feature information can also be understood as the initial feature information of the first sample, and the first feature information includes the initial feature information of each first information and the initial feature information of the first time interval corresponding to each first information. After the execution device inputs the first feature information into the first machine learning model, the first machine learning model performs feature update and feature processing on the first feature information, and obtains the prediction information corresponding to the second time period output by the first machine learning model.

[0117] For example, in an application scenario, when the first information is the alarm information of the network device, the prediction information corresponding to the second time period can indicate whether the first fault will occur in the second time period, and the time point when the first information is generated is the time point when the alarm information occurs. In the embodiment of the present application, a specific application scenario of the present application is provided, which improves the degree of integration between the present solution and the specific application scenario.

[0118] Exemplarily, the aforementioned network equipment may be specifically manifested as a base station, a network equipment in a core network, or a network equipment in other forms, etc. The form of the specific network equipment may be flexibly determined in combination with the actual application scenario and is not limited here.

[0119] Exemplarily, "alarm type" can also be understood as "alarm name", and alarm types include but are not limited to radio frequency unit Zhu Bo alarm, optical network unit (ONU) signal loss, or low temperature, etc. It should be noted that the examples here are only for the convenience of understanding the concept of alarm type and are not used to limit this solution. The first fault may include any one or more of the following faults: base station decommissioning fault, cell decommissioning fault, optical power fault, single board temperature fault, combined wave link interruption fault or optical path interruption fault, etc. The specific manifestation of the "first fault" can be determined in combination with the actual application scenario, and is not limited here.

[0120] For another example, in another application scenario, when the first information is the user's physiological information, the prediction information corresponding to the second time period can indicate the user's motion state in the second time period, and the time point when the first information is generated is the time point when the physiological information occurs. In the embodiment of the present application, another specific application scenario of the present application is provided to improve the implementation flexibility of the present solution.

[0121] Exemplarily, the user's physiological information includes but is not limited to: heart rate, arm swing amplitude or other physiological information; the user's motion state includes but is not limited to: walking state, running state, going upstairs state, going downstairs state or other motion state, etc. The examples given here are only for the convenience of understanding this solution. The specific types of physiological information and the specific motion states to be used can be determined in combination with the actual application scenarios, and are not limited in the embodiments of this application.

[0122] It should be noted that the "first information" can also be expressed as other information, and the "prediction information corresponding to the second time period" can also indicate other information, etc. The specific information can be determined in combination with the actual application scenario, and is not listed here in an exhaustive manner.

[0123] Exemplarily, the second time period may intersect with the first time period, that is, there is an overlapping time period between the first time period and the second time period; or, there may be no intersection between the second time period and the first time period, that is, there is no overlapping time period between the first time period and the second time period. When there is no intersection between the second time period and the first time period, in one case, the second time period and the first time period may be adjacent, that is, the end time point of the first time period may be the start time point of the second time period; or, in another case, there is an interval between the first time period and the second time period, that is, after the end time point of the first time period, it may be a period of time before the start time point of the second time period; it should be understood that the "relationship between the first time period and the second time period" can be flexibly set in combination with the actual application scenario, and is not limited in the embodiments of the present application.

[0124] In an embodiment of the present application, the input of the machine learning model includes not only each first information generated within the first time period, but also first time information, where the first time information indicates the relative time interval between the time point at which each first information is generated and a preset reference time point. When the first information generated in two different first time periods (for example, first time period 1 and first time period 2) is the same, the first time information corresponding to the first information generated in first time period 1 may be different from the first time information corresponding to the first information generated in first time period 2. That is, even if the first information generated in two different first time periods is the same, the two different first time periods can be distinguished with the help of the aforementioned preset reference time point, thereby facilitating the machine learning model to better distinguish different situations, which is beneficial to improving the accuracy of the prediction information output by the machine learning model. In addition, during the training stage of the machine learning model, when the machine learning model encounters a "junk sample pair", although the multiple first information generated in the historical data of M time units included in training sample A and training sample B are the same, since the training samples are collected using a sliding window, the relative time interval between the time point at which the multiple first information is generated in training sample A and the first time point is different from the relative time interval between the multiple first information generated in training sample B and the first time point. After the relative time interval is added to the input of the machine learning model, training sample A and training sample B are different for the machine learning model. When the true values ​​included in training sample A and training sample B are different, the machine learning model can also learn the mapping relationship between the historical data and the true value in the first time period, which is beneficial to improve the accuracy of the prediction information output by the machine learning model.

[0125] In combination with the above description, the detailed implementation process of the training phase of the first machine learning model and the reasoning phase of the first machine learning model are introduced below.

[0126] 1. Training Phase

[0127] For details, please refer to Figure 4 , Figure 4 Another schematic diagram of the data processing method provided in the embodiment of the present application is as follows: Figure 4 As shown, the data processing method provided by the present application may include:

[0128] 401. Obtain a first training sample, where the first training sample includes at least one first information generated in a third time period.

[0129] In an embodiment of the present application, a training sample set may be deployed in the training device, and the training sample set includes multiple training samples and a ground truth corresponding to each training sample. Exemplarily, each of the multiple training samples may include at least one first information generated within M time units, where M is an integer greater than or equal to 1.

[0130] The “true value corresponding to the training sample” can also be called the “expected information corresponding to the training sample”, or can also be called the “label corresponding to the training sample”, etc.

[0131] For example, in an application scenario, the first information may be expressed as an alarm information, and the alarm information may carry an alarm type. The true value corresponding to each training sample may indicate whether the first fault will occur within K time units, where K is an integer greater than or equal to 1. K time units are located after M time units, that is, K time units are later than M time units. For examples of "alarm type" and "first fault", please refer to the above Figure 3 The description in the corresponding embodiment will not be repeated here.

[0132] In this application scenario, the time units in "M time units" and "K time units" can be expressed as days, hours or other time lengths, etc.; for example, the length of M time units can be 7 days, 5 days, 3 days or other durations, etc., and the length of K time units can be 3 days, 2 days, 1 day or other durations, etc. The specific lengths of "M time units" and "K time units" can be determined in combination with the actual application scenario and are not limited here.

[0133] For the convenience of description, in this application, any training sample in the training data set is referred to as a "first training sample", and the first training sample includes at least one first information generated within a third time period (that is, an example of M time units), and the true value corresponding to the first training sample indicates whether a first fault will occur within a fourth time period (that is, K time units).

[0134] Optionally, in this application scenario, when the true value corresponding to the first training sample indicates that the first fault will occur within the fourth time period, the first training sample is a positive training sample; when the true value corresponding to the first training sample indicates that the first fault will not occur within the fourth time period, the first training sample is a negative training sample.

[0135] Optionally, the embodiment of the present application also provides a "process for obtaining multiple training samples in a training sample set". Exemplarily, the training samples in the training sample set can be obtained based on historical alarm data collected by at least one communication device, and the historical alarm data may include multiple alarm information collected by each communication device and the generation time of each alarm information. It should be noted that the step of "constructing training samples based on historical alarm data" can be performed by the training device; or, it can also be performed by other devices other than the training device, and then the other devices send the training data set to the training device. The specific implementation method is not limited here.

[0136] For the convenience of description, any one of the at least one communication device is referred to as a "target communication device". Optionally, for the historical alarm data collected from the target communication device, a training sample may be constructed in a sliding window manner.

[0137] Exemplarily, since the time length of each training sample is M time units, the true value corresponding to each training sample is obtained based on the historical alarm data of K time units after M time units, then the window length corresponding to each training sample can be M+K time units, and the window can be slid according to a preset sliding time interval (for example, T time units) to obtain multiple second samples from the historical alarm data collected by the target communication device, and each of the multiple second samples includes the alarm information generated within M+K time units (that is, historical alarm information); the above steps are performed on the historical alarm data collected by each communication device in at least one communication device, and multiple second samples can be obtained based on all historical alarm data collected by at least one communication device.

[0138] For the convenience of description, any second sample among the multiple second samples is referred to as a "target sample", and the training sample corresponding to the target sample is referred to as a target training sample. The true value corresponding to the target training sample can be determined based on the alarm information within the K time units included in the target sample. Exemplarily, if there is a first alarm type that can directly indicate a first fault in all the alarm information within the K time units included in the target sample, it is determined that the true value corresponding to the target training sample indicates that the first fault will occur within K time units; if there is no first alarm type that can directly indicate the first fault in all the alarm information within the K time units included in the target sample, it is determined that the true value corresponding to the target training sample indicates that the first fault will not occur within K time units.

[0139] “The first alarm type that can directly indicate the first fault” can be understood as when it is determined that the first alarm type exists within K time units, it means that the first fault will occur within K time units; the “first alarm type” specifically refers to which alarm types can be determined in combination with the type of the first fault or other factors. For example, when the first fault is a base station decommissioning fault, the first alarm type may include base station out of service (gNodeB Out of Service), communication service company error (CSL Fault), or other alarm types; for another example, when the first fault is a rectifier fault, the first alarm type may include rectifier module fault, rectifier fault (Rectifier Fault), or other types of alarm types. It should be noted that the examples given here are only for the convenience of understanding the indication relationship between the “first alarm type” and the “first fault”, and are not used to limit this solution.

[0140] Based on the historical alarm data within the M time units included in the target sample, it is possible to determine which alarm types are carried in at least one first information included in the target training sample. In one case, it is also possible to determine N alarm types related to the first fault based on multiple positive second samples; illustratively, a frequent item mining algorithm can be used to determine the N alarm types most relevant to the first fault based on multiple positive second samples. Among them, if there is a first alarm type that can directly indicate the first fault in all the alarm information within the K time units included in the second sample, the second sample can be marked as a positive sample; if there is no first alarm type that can directly indicate the first fault in all the alarm information within the K time units included in the second sample, the second sample can be marked as a negative sample.

[0141] After determining N alarm types, the historical alarm data within M time units included in the target sample are screened to obtain a target training sample. All alarm types carried by at least one alarm information in the target training sample are included in the aforementioned N alarm types, and other alarm types outside the aforementioned N alarm types have been eliminated.

[0142] In another case, all alarm types carried in the historical alarm data within M time units may be determined as the alarm types in the target training samples, that is, there is no need to eliminate any alarm types in the historical alarm data within M time units.

[0143] Optionally, since the number of alarm information carried in different training samples may be different, the number of alarm information included in each training sample can also be set to a first value, and the target training sample is updated based on the aforementioned first value to obtain an updated target training sample included in the training sample set. In one case, if the number of alarm information included in the target training sample is greater than the first value, at least one alarm information in the target training sample can be deleted so that the number of alarm information included in the updated target training sample is equal to the first value. Optionally, for example, the number of the first value is S, and S is an integer greater than or equal to 1. According to the termination time point of the target training sample and the generation time of each alarm information in the target training sample, the S alarm information whose generation time is closest to the termination time point can be determined, and the other alarm information in the target training sample can be deleted.

[0144] Exemplarily, the value of S is 5, the end time point of the target training sample is time point 1, the target training sample includes 8 alarm information, and the generation time of the 8 alarm information is from early to late: time point 2, time point 3, time point 4, time point 5, time point 6, time point 7, time point 8 and time point 9. The relative time interval between time point 2 and time point 1 is the longest, and the relative time interval between time point 9 and time point 1 is the shortest. Then, the alarm information generated at time point 5, time point 6, time point 7, time point 8 and time point 9 can be retained, and the alarm information generated at time point 2, time point 3 and time point 4 can be deleted. That is, the updated target training sample includes the alarm information generated at time point 5, time point 6, time point 7, time point 8 and time point 9. It should be understood that the examples here are only for the convenience of understanding of this scheme and are not used to limit this scheme.

[0145] In another case, if the amount of warning information included in the target training sample is less than the first value, meaningless warning information may be added to the target training sample so that the amount of warning information included in the updated target training sample is equal to the first value.

[0146] It should be noted that the above step of "obtaining a target training sample in a training sample set based on a target sample" can be adopted to perform the above operation on each second sample, so that multiple training samples in a training sample set can be obtained according to multiple second samples.

[0147] In another application scenario, the first information includes the user's physiological information, and the true value corresponding to each training sample can indicate the user's motion state within K time units. For examples of "user's physiological information" and "motion state", please refer to the above Figure 3 The description in the corresponding embodiment is not repeated here.

[0148] In this application scenario, the time units in "M time units" and "K time units" can be expressed as minutes, seconds or other time lengths, etc.; for example, the length of M time units can be 15 minutes, 10 minutes or 5 minutes, etc., and the length of the fourth time period can be 3 minutes, 2 minutes or 1 minute, etc. The specific length of the third time period and the length of the fourth time period can be determined in combination with the actual application scenario, and are not limited in this application.

[0149] Optionally, the first training sample may further include second time information, the second time information indicating the time point when each first information is generated; illustratively, the second time information may include a timestamp of each first information, the timestamp of each first information indicating the time point when each first information is generated.

[0150] Optionally, the first training sample may further include identification information of each first information in the at least one first information; for example, the identification information of each first information may be expressed as an identification (ID) of each first information.

[0151] 402. Obtain first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point.

[0152] In the embodiment of the present application, step 402 is an optional step. For example, the first time point may be any time point with a fixed relative time interval from the third time period; the first time point may be within the third time period or outside the third time period.

[0153] Optionally, the first time point may be determined based on the starting time point or the ending time point of the third time period. For example, the first time point may be the starting time point of the third time period. For another example, the first time point may be the ending time point of the third time period. For another example, the first time point may be a time point located in the middle of the third time period, and the aforementioned time point located in the middle is determined based on the starting time point and the ending time point of the third time period. For another example, the first time point is 2 minutes earlier than the starting time point of the third time period. For another example, the first time point is 2 minutes later than the starting time point of the third time period. For another example, the first time point is 3 minutes earlier than the ending time point of the third time period. For another example, the first time point is 5 minutes later than the ending time point of the third time period, etc. It should be understood that the specific setting of the "first time point" can be flexibly determined in combination with the actual application scenario, and the example of the "first time point" here is only for the convenience of understanding this solution, and is not used to limit this solution.

[0154] In an embodiment of the present application, a first time point is determined based on the starting time point or the ending time point of the third time period, and a scheme for determining the first time point is provided, thereby reducing the difficulty of implementing the scheme; optionally, when the first time point adopts the starting time point or the ending time point of the third time period, the difficulty of the process of "determining the first time point" is further reduced.

[0155] In an embodiment of the present application, after acquiring at least one first information included in the first training sample, the training device can determine the time point when each first information is generated, and then determine the relative time interval between the time point when each first information is generated and the first time point, that is, obtain the first time interval between the time point when each first information is generated and the first time point.

[0156] Exemplarily, in one implementation, the first training sample includes not only at least one first information but also second time information. In step 402, the training device can determine the time point at which each first information is generated based on the second time information in the first training sample. In another implementation, if the first training sample does not include the second time information, in step 402, the training device can additionally obtain the second time information, and then determine the time point at which each first information is generated based on the obtained second time information. It should be understood that the specific expression of the "second time information" can be found in the description in step 401, which will not be repeated here.

[0157] For a more intuitive understanding of the first-time information, please refer to Figure 5 , Figure 5 A schematic diagram of the first time information provided in an embodiment of the present application. Figure 5 The first time point (i.e. Figure 5 T n) is the end time point of the third time period, and the first information is an alarm information, for example, Figure 5 The circle in the figure represents that the alarm information 1 (ie, an example of the first information) is generated at T1. The first time information includes T n A first time interval 1 from T1; Figure 5 The five-pointed star in the figure represents that the alarm information 2 (ie, another example of the first information) is generated at T2. The first time information also includes T n The first time interval 2 between T2 and T3; it should be understood that Figure 5 The meaning of "first time information" is only understood by taking two alarm messages occurring in the third time period as an example, and is not used to limit this solution.

[0158] 403. Obtain first characteristic information, where the first characteristic information includes characteristic information of the first information and characteristic information of a first time interval corresponding to the first information.

[0159] In an embodiment of the present application, after acquiring the first training sample and the first time information, the training device can obtain the first feature information; exemplarily, the feature information of each first information in the first feature information can also be understood as "the initial feature information of each first information", and the feature information of each first time interval in the first feature information can also be understood as "the initial feature information of each first time interval".

[0160] A specific implementation method for obtaining the first characteristic information for the training device. In an application scenario, if each first information carries an alarm type, then in an implementation method, the second information may be deployed in the training device, wherein the second information indicates the characteristic information corresponding to each of the multiple first information, and the second information further indicates the characteristic information corresponding to each of the multiple time intervals, and the first information is updated during the training of the first machine learning model. Exemplarily, if the first information is an alarm message, and the alarm message carries an alarm type, then the second information may include the characteristic information corresponding to each of the multiple alarm types. Optionally, the alarm type carried in the meaningless alarm message is meaningless, and the aforementioned multiple alarm types may include meaningless alarm types, that is, the second information may include the characteristic information of the meaningless alarm type.

[0161] Step 403 may include: after the training device determines at least one first information included in the first training sample and at least one first time interval included in the first time information, the training device can obtain the initial characteristic information of the alarm type carried by each first information and the initial characteristic information of each first time interval from the second information.

[0162] In another implementation, step 403 may include: the training device may vectorize the alarm type carried in each first information to obtain initial feature information of each first information; and vectorize each first time interval in each first time information to obtain initial feature information of each first time interval. Optionally, the aforementioned vectorization step may be performed by a second machine learning model.

[0163] Exemplarily, the training device may perform word embedding on the alarm type carried in each first message to obtain initial feature information of each first message; and perform word embedding on each first time interval to obtain initial feature information of each first time interval.

[0164] In another application scenario, if each first information includes physiological information of the user, and the first training sample includes multiple physiological information generated by the user in a third time period, step 403 may include: the training device vectorizes each first information in the first training sample to obtain initial feature information of each first information; and vectorizes each first time interval in the first time information to obtain initial feature information of each first time interval. Optionally, the aforementioned vectorization step may be performed by a second machine learning model.

[0165] In step 403, optionally, the training device may also combine the initial feature information of each first information with the initial feature information of the first time interval corresponding to each first information to obtain a second feature information corresponding to a first information, then the first feature information may include at least one second feature information corresponding to at least one first information, and each second feature information includes the initial feature information of a first information and the initial feature information of the first time interval corresponding to the aforementioned first information. Further, for the convenience of distinction, any second feature information in at least one second feature information is referred to as "target second feature information", and "target second feature information corresponds to target information in at least one first information" can be understood as the target second feature information includes the feature information of the target information.

[0166] Exemplarily, the first training sample includes 5 first information, namely, first information 1, first information 2, first information 3, first information 4 and first information 5, the first time information includes the first time interval 1 corresponding to the first information 1, the first time interval 2 corresponding to the first information 2, the first time interval 3 corresponding to the first information 3, the first time interval 4 corresponding to the first information 4 and the first time interval 5 corresponding to the first information 5, and the first feature information may include 5 second feature information. Among them, the second feature information 1 is obtained by combining the initial feature information of the first information 1 and the initial feature information of the first time interval 1, the second feature information 2 is obtained by combining the initial feature information of the first information 2 and the initial feature information of the first time interval 2, the second feature information 3 is obtained by combining the initial feature information of the first information 3 and the initial feature information of the first time interval 3, the second feature information 4 is obtained by combining the initial feature information of the first information 4 and the initial feature information of the first time interval 4, and the second feature information 5 is obtained by combining the initial feature information of the first information 5 and the initial feature information of the first time interval 5. It should be understood that the examples here are only for the convenience of understanding of this solution and are not used to limit this solution.

[0167] 404. Input the first feature information into the first machine learning model to obtain prediction information corresponding to a fourth time period output by the first machine learning model, where the fourth time period is located after the third time period.

[0168] In an embodiment of the present application, after obtaining the first characteristic information, the training device may input the first characteristic information into the first machine learning model, process the first characteristic information through the first machine learning model, and obtain the prediction information corresponding to the fourth time period output by the first machine learning model. For example, if the first information is an alarm information, the prediction information corresponding to the fourth time period may indicate whether the first fault will occur in the second time period; for another example, if the first information includes the user's physiological information, the prediction information corresponding to the fourth time period may indicate the user's motion state in the fourth time period, etc. The meanings of "alarm information", "first fault", "user's physiological information" and "motion state" can all be referred to the above description, and will not be repeated here.

[0169] The fourth time period is located after the third time period. The length of the third time period may be M time units, and the length of the fourth time period may be K time units. The meanings of “M time units” and “K time units” may also be found in the above description and will not be elaborated here.

[0170] Exemplarily, the training device processes the first feature information through the first machine learning model, which may include: the training device updates the first feature information through the first machine learning model to obtain updated first feature information; wherein the feature information generated in the process of "the training device updates the first feature information through the first machine learning model" can be understood as the feature information of the first training sample. The training device performs feature processing on the updated first feature information through the first machine learning model to obtain prediction information corresponding to the fourth time period.

[0171] The first feature information includes at least one second feature information corresponding one-to-one to at least one first information. The meaning of the aforementioned "second feature information" can be found in the description in step 403 and will not be elaborated here; "updating the first feature information through the first machine learning model" can be understood as updating the features of each second feature information through the first machine learning model.

[0172] Optionally, the first machine learning model may include one or more first feature update modules. In the process of updating the first feature information through the first feature update module in the first machine learning model, the method also includes: the training device divides at least one second feature information into at least two groups based on each of at least two grouping bases, wherein the target grouping basis is any one of the at least two grouping bases, and based on the target grouping basis, the at least one second feature information is divided into at least two target groups.

[0173] The training device updates each second feature information in the same target group in at least two target groups through the first feature update module, based on the second feature information in the same target group, to obtain an updated feature of each second feature information. Optionally, the training device updates each second feature information in the same target group based on the self-attention mechanism, based on the second feature information in the same target group in at least two target groups, to obtain an updated feature of each second feature information.

[0174] Since the training device updates the second feature information based on at least two grouping criteria, the training device repeats the above operation at least twice, and can obtain at least two updated features of each second feature information; the training device combines at least two updated features of each second feature information through the first feature update module, and obtains each updated second feature information, thereby completing an update of the second feature information using the first feature update module.

[0175] Exemplarily, the first feature information includes 6 second feature information, namely, second feature information A, second feature information B, second feature information C, second feature information D, second feature information E, and second feature information F. The training device groups the aforementioned 6 second feature information based on each of the 2 grouping criteria, and divides the aforementioned 6 second feature information into 2 groups using the first grouping basis, the first of the 2 groups includes the second feature information A, the second feature information B, and the second feature information C, and the second of the 2 groups includes the second feature information D, the second feature information E, and the second feature information F. The training device updates the second feature information A, the second feature information B and the second feature information C based on the self-attention mechanism, and obtains the updated feature 1 of the second feature information A, the updated feature 1 of the second feature information B and the updated feature 1 of the second feature information C; according to the second feature information D, the second feature information E and the second feature information F, the training device updates the second feature information D, the second feature information E and the second feature information F based on the self-attention mechanism, and obtains the updated feature 1 of the second feature information D, the updated feature 1 of the second feature information E and the updated feature 1 of the second feature information F.

[0176] The training device divides the above-mentioned 6 second feature information into 3 groups using the second grouping basis, the first group of the 3 groups includes the second feature information A and the second feature information D, the second group of the 3 groups includes the second feature information B, the second feature information C and the second feature information E, and the third group of the 3 groups includes the second feature information F; the training device updates the second feature information A and the second feature information D based on the self-attention mechanism according to the second feature information A and the second feature information D, and can obtain the updated feature 2 of the second feature information A and the updated feature 2 of the second feature information D; according to the second feature information B, the second feature information C and the second feature information E, the second feature information B, the second feature information C and the second feature information E are updated based on the self-attention mechanism, and the updated feature 2 of the second feature information B, the updated feature 2 of the second feature information C and the updated feature 2 of the second feature information E can be obtained; according to the second feature information F, the second feature information F is updated based on the self-attention mechanism to obtain the updated feature 2 of the second feature information F.

[0177] The training device combines the updated feature 1 of the second feature information A with the updated feature 2 of the second feature information A to obtain the updated second feature information A; combines the updated feature 1 of the second feature information B with the updated feature 2 of the second feature information B to obtain the updated second feature information B; combines the updated feature 1 of the second feature information C with the updated feature 2 of the second feature information C to obtain the updated second feature information C; combines the updated feature 1 of the second feature information D with the updated feature 2 of the second feature information D to obtain the updated second feature information D; combines the updated feature 1 of the second feature information E with the updated feature 2 of the second feature information E to obtain the updated second feature information E; combines the updated feature 1 of the second feature information F with the updated feature 2 of the second feature information F to obtain the updated second feature information F, thereby completing a feature update from the second feature information A to the second feature information F. It should be understood that the examples given here are only for the convenience of understanding this scheme and are not used to limit this scheme.

[0178] It should be noted that if there are multiple first feature update modules in the first machine learning model, the multiple first feature update modules can be used to perform multiple feature updates on the first feature information; for example, the multiple first feature update modules can be serial, and the training device can use a first feature update module to perform feature update on each second feature information in the first feature information to obtain each updated second feature information; then use another first feature update module to perform feature update on each updated second feature information again to obtain the updated second feature information again, and the training device repeats the aforementioned steps multiple times to complete the feature update of the second feature information using each first feature update module.

[0179] In an embodiment of the present application, multiple second feature information are grouped based on at least two different perspectives, and then the second feature information in the same group after grouping is updated, so that the feature update process of the first feature information can be very sufficient, which is conducive to the first machine learning model learning rich information, thereby facilitating improving the accuracy of the prediction information output by the first machine learning model.

[0180] Optionally, if the first information is an alarm message, and the alarm message carries an alarm type, each second characteristic information includes characteristic information of an alarm message and characteristic information of a first time interval corresponding to the aforementioned alarm message. Since each second characteristic information corresponds one-to-one with a first information (i.e., an alarm message), illustratively, the at least two grouping criteria include any at least two of the following: the generation time of the first information (i.e., the alarm message), whether the alarm type included in the first information (i.e., the alarm message) can directly indicate the first fault, the alarm type in the first information (i.e., the alarm message), random grouping, or other grouping criteria, etc. In the embodiment of the present application, a variety of specific grouping criteria are provided to improve the implementation flexibility of the present solution; in addition, the grouping criteria include the generation time of the alarm message, whether the alarm message can directly indicate the first fault, and the type of alarm type carried by the alarm message. These grouping criteria are very suitable for the application scenario of "the first information is an alarm message", which is more conducive to the first machine learning model to obtain sufficient information and to improve the accuracy of the prediction information output by the first machine learning model.

[0181] Exemplarily, in one case, if the grouping basis is the generation time of the alarm information, the third time period is divided into at least two sub-time periods, and the training device can divide at least one second characteristic information into at least two first groups corresponding to the at least two sub-time periods, each first group including the second characteristic information corresponding to all the first information generated in one of the at least two sub-time periods.

[0182] For a more intuitive understanding of this solution, please refer to Figure 6 , Figure 6 A schematic diagram of grouping at least one second feature information provided in an embodiment of the present application. Figure 6 In the figure, Alarm A represents alarm information A, Alarm B represents alarm information B, and Alarm C represents alarm information C. Figure 6 Taking the length of the third time period as 7 days as an example, 7 days are divided into 3 sub-time periods, which are the first 3 days, the middle 2 days and the last 2 days within the 7 days. The multiple second characteristic information corresponding to the multiple alarm information generated within the 7 days are divided into 3 first groups, which include the first group 1, the first group 2 and the first group 3. Among them, the multiple second characteristic information in the first group 1 corresponds to the alarm information generated within the first 3 days, the multiple second characteristic information in the first group 2 corresponds to the alarm information generated within the middle 2 days, and the multiple second characteristic information in the first group 3 corresponds to the alarm information generated within the last 2 days. It should be understood that Figure 6 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0183] The training device updates each second feature information in each first group according to the second feature information in each first group of at least two first groups to obtain an updated feature of each second feature information. Optionally, the training device updates each second feature information in each first group based on a self-attention mechanism according to the second feature information in each first group.

[0184] To further understand this solution, Figure 6 For example, multiple second feature information corresponding to multiple alarm information generated within 7 days are divided into three first groups, and the three first groups include first group 1, first group 2 and first group 3. Among them, the training device uses the multiple second feature information in the first group 1 to update the multiple second feature information in the first group 1 based on the self-attention mechanism. The training device uses the multiple second feature information in the first group 2 to update the multiple second feature information in the first group 2 based on the self-attention mechanism. The training device uses the multiple second feature information in the first group 3 to update the multiple second feature information in the first group 3 based on the self-attention mechanism; through the aforementioned steps, the training device can obtain an updated feature of each second feature information. It should be understood that the examples given here are only for the convenience of understanding this solution and are not used to limit this solution.

[0185] In another case, if the grouping basis is whether the alarm type included in the alarm information can directly indicate the first fault, the training device can divide at least one second characteristic information into a second group and a third group, the second group includes second characteristic information corresponding to all first information in at least one first information that can directly indicate the first fault, and the third group includes second characteristic information corresponding to all first information in at least one first information other than the second group.

[0186] For a more intuitive understanding of this solution, please refer to Figure 7 , Figure 7 Another schematic diagram of grouping at least one second feature information provided in an embodiment of the present application, Figure 7 In the example, Alarm A represents alarm information A, Alarm B represents alarm information B, and Alarm C represents alarm information C. Figure 7 As shown, Alarm B can directly indicate the first fault, then the second group includes the second characteristic information corresponding to all Alarm B generated in the third time period, and the third group includes the second characteristic information corresponding to all alarm information other than Alarm B in all alarm information generated in the third time period. It should be understood that Figure 7 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0187] The training device updates each second feature information in the second group according to the second feature information in the second group, and updates each second feature information in the third group according to the second feature information in the third group, thereby obtaining another updated feature of each second feature information. Optionally, the training device updates each second feature information in the second group based on the self-attention mechanism according to the second feature information in the second group; and updates each second feature information in the third group based on the self-attention mechanism according to the second feature information in the third group.

[0188] In another case, if the grouping basis is the alarm type in the alarm information, the training device can divide at least one second characteristic information into at least two fourth groups, and the alarm type carried in the first information (i.e., the alarm information) corresponding to all the second characteristic information in each fourth group is the same.

[0189] For a more intuitive understanding of this solution, please refer to Figure 8 , Figure 8 Another schematic diagram of grouping at least one second feature information provided in an embodiment of the present application, Figure 8 In the above, Alarm A represents alarm information A, Alarm B represents alarm information B, and Alarm C represents alarm information C. The multiple second characteristic information corresponding to the multiple alarm information generated in the third time period are divided into three fourth groups, the three fourth groups include fourth group 1, fourth group 2, and fourth group 3, fourth group 1 includes the second characteristic information corresponding to all Alarm A, fourth group 2 includes the second characteristic information corresponding to all Alarm B, and fourth group 3 includes the second characteristic information corresponding to all Alarm C. It should be understood that Figure 8 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0190] The training device updates each second feature information in each fourth group according to the second feature information in each fourth group of at least two fourth groups to obtain an updated feature of each second feature information. Optionally, the training device updates each second feature information in each fourth group based on a self-attention mechanism according to the second feature information in each fourth group.

[0191] To further understand this solution, Figure 8For example, the alarm types in all the alarm information generated within 7 days include alarm type A, alarm type B, and alarm type C. The multiple second characteristic information corresponding to the multiple alarm information generated in the third time period are divided into three fourth groups, and the three fourth groups include fourth group 1, fourth group 2, and fourth group 3. Among them, the alarm types in the alarm information corresponding to the three second characteristic information in fourth group 1 are all alarm type A, and the training device uses the multiple second characteristic information in fourth group 1 to update the multiple second characteristic information in fourth group 1 based on the self-attention mechanism. The alarm types in the alarm information corresponding to the two second characteristic information in fourth group 2 are all alarm type B, and the training device uses the multiple second characteristic information in fourth group 2 to update the multiple second characteristic information in fourth group 2 based on the self-attention mechanism. The alarm types in the alarm information corresponding to the three second feature information in the fourth group 3 are all alarm type C. The training device uses the multiple second feature information in the fourth group 3 to update the multiple second feature information in the fourth group 3 based on the self-attention mechanism; through the above steps, the training device can obtain an updated feature of each second feature information. It should be understood that the examples given here are only for the convenience of understanding this scheme and are not used to limit this scheme.

[0192] In an embodiment of the present application, specific implementation steps for updating the features of the first feature information using the first machine learning model are provided, which is conducive to reducing the difficulty of implementing the present solution.

[0193] To more intuitively understand the process of updating the features of each second feature information using the first feature update module in the first machine learning model, please refer to Fig. 9 , Fig. 9 A schematic diagram of updating features through a first machine learning model provided in an embodiment of the present application. Fig. 9 As shown, in the process of updating the first feature information through the first machine learning model, after the first feature information is updated L times through the first machine learning model, the updated first feature information is obtained, and L is an integer greater than or equal to 1.

[0194] Exemplarily, in the process of updating the first feature information each time by the first machine learning model, the input feature information (i.e. Fig. 9 The input Embedding in the input Embedding may be the first feature information, or may be the updated first feature information obtained after the feature of the first feature information is last updated.

[0195] Then, the input Embedding is updated by the first feature updating module. Specifically, the input Embedding can be divided into at least two groups by using each of the three grouping criteria. Since the input Embedding includes a plurality of second feature information corresponding to a plurality of alarm information generated in the third time period, such as Fig. 9 As shown, in (a) the grouping basis, multiple alarm information can be first divided into at least two groups corresponding to at least two sub-time periods based on the time point when each alarm information is generated, so that multiple second feature information corresponding to multiple alarm information generated in the third time period can be divided into at least two first groups, each first group including second feature information corresponding to all alarm information generated in one of the at least two sub-time periods; then based on the self-attention mechanism, the second feature information in each first group is updated respectively to obtain the updated feature 1 of each second feature information.

[0196] In (b) the grouping basis, all alarm information can be divided into alarm information related to the label and alarm information not related to the label. "Alarm information related to the label" can be understood as "alarm information that can directly indicate the first fault", and "alarm information not related to the label" can be understood as "alarm information that cannot directly indicate the first fault". Then, multiple second feature information corresponding to multiple alarm information generated in the third time period can be divided into a second group and a third group, the second group includes all second feature information corresponding to the alarm information that can directly indicate the first fault, and the third group includes all second feature information corresponding to the alarm information except the second group in the multiple alarm information; then based on the self-attention mechanism, the second feature information in the second group and the third group are updated respectively to obtain the updated feature 2 of each second feature information.

[0197] In (c) the grouping basis, the multiple alarm information generated in the third time period can be grouped according to the alarm type of each alarm information, and the multiple second feature information corresponding one by one to the multiple alarm information generated in the third time period can be divided into at least two fourth groups, and the alarm types carried in the alarm information corresponding to all the second feature information in each fourth group are the same; then based on the self-attention mechanism, the second feature information in each fourth group is updated respectively to obtain the updated feature 3 of each second feature information.

[0198] The updated feature 1, the updated feature 2 and the updated feature 3 of each second feature information are aggregated to obtain each updated second feature information, that is, the updated first feature information generated by the first feature update module; in this round, the updated first feature information generated by the first feature update module is also subjected to residual accumulation (Add) & normalization (normalization, Norm) operation and feed forward (Feed Forward) operation to obtain the updated first feature information generated in this round.

[0199] Repeat the above operation L times to obtain the updated first feature information generated by the entire feature updating module of the first machine learning model, and then perform feature processing on the updated first feature information generated by the entire feature updating module through the feature processing module of the first machine learning model to obtain the prediction information output by the entire first machine learning model. It should be understood that Fig. 9 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0200] 405. Train the first machine learning model according to a first loss function, where the first loss function indicates a similarity between the prediction information corresponding to the fourth time period and the true value corresponding to the first training sample.

[0201] In an embodiment of the present application, the training device can generate a function value of a first loss function based on the similarity between the prediction information corresponding to the fourth time period and the true value corresponding to the first training sample; based on the function value of the first loss function, the parameters of the first machine learning model are gradient updated based on the back propagation algorithm to achieve one training of the first machine learning model.

[0202] Optionally, if the first feature information in step 403 is generated based on the second information, then in step 405, the training device performs a gradient update on the parameters of the first machine learning model based on the back propagation algorithm, which may include: the training device performs a gradient update on the parameters of the first machine learning model based on the back propagation algorithm, and performs a gradient update on the second information.

[0203] Optionally, if the first feature information in step 403 is generated based on the second machine learning model, then in step 405, the training device performs a gradient update on the parameters of the first machine learning model based on the back propagation algorithm, which may include: the training device performs a gradient update on the parameters of the first machine learning model and the second machine learning model based on the back propagation algorithm.

[0204] Optionally, in the case where the first information is expressed as warning information, the category of each training sample can be a positive training sample or a negative training sample, and step 405 may include: the training device generates a total loss function value according to the function value of the first loss function and the function value of the second loss function, and then according to the total loss function value, performs a gradient update on the parameters of the first machine learning model based on the back propagation algorithm.

[0205] Optionally, the training device performs a gradient update on the parameters of the first machine learning model based on the back propagation algorithm according to the total loss function value, and performs a gradient update on the second information. Alternatively, the training device performs a gradient update on the parameters of the first machine learning model and the parameters of the second machine learning model based on the back propagation algorithm according to the total loss function value.

[0206] In an embodiment of the present application, second information is deployed in the first device, and the second information includes characteristic information of each alarm type and characteristic information of each time interval. The initial characteristic information corresponding to the first sample (that is, the first characteristic information) can be quickly acquired based on the second information, and the second information will also be updated during the training phase of the machine learning model, which is conducive to obtaining more accurate initial characteristic information of the first sample.

[0207] Exemplarily, the second loss function indicates the similarity between feature information of different positive training samples, that is, the second loss function is obtained based on the feature information of the positive training samples. The goal of training using the second loss function includes improving the similarity between feature information of different positive training samples. Multiple training samples can be used in a training operation of the first machine learning model. For example, the aforementioned multiple training samples may include a batch of training samples, and the multiple training samples may include one or more positive training samples. The concept of "positive training sample" can be found in the description of step 401 above and will not be elaborated here.

[0208] The training device can obtain the feature information of each positive training sample from step 404. In step 405, the training device can generate the function value of the second loss function based on the feature information of at least one positive training sample. The second loss function indicates the similarity between the feature information of different positive training samples. The meaning of "feature information of positive training samples" is similar to that of "feature information of first training samples". Please refer to it for understanding and will not be repeated here.

[0209] Optionally, the second loss function is obtained based on the feature information of the positive training samples and the feature information of the negative training samples. The second loss function also indicates the similarity between the feature information of different negative training samples. The purpose of training using the second loss function also includes improving the similarity between the feature information of different negative training samples.

[0210] Optionally, the second loss function is obtained based on the feature information of the positive training samples and the feature information of the negative training samples. The second loss function also indicates the similarity between the feature information of the positive training samples and the feature information of the negative training samples. The goal of training using the second loss function also includes reducing the similarity between the feature information of the positive training samples and the feature information of the negative training samples.

[0211] Exemplarily, a batch of training samples may include one or more negative training samples. The concept of "negative training sample" can be found in the description in step 401 above. The meaning of "characteristic information of negative training sample" is similar to that of "characteristic information of the first training sample", which can be referred to for understanding and will not be elaborated here.

[0212] Then, in step 405, the training device may generate a function value of a second loss function according to feature information of at least one positive training sample and feature information of at least one negative training sample.

[0213] For a more intuitive understanding of this solution, please refer to Fig.10 , Fig.10 A schematic diagram of multiple information indicated by the second loss function provided in an embodiment of the present application, Fig.10 In the example, five training samples are used in a training of the first machine learning model, including two negative training samples and three positive training samples. After obtaining the first feature information of each training sample in the five training samples, the first feature information of each training sample is updated through the first machine learning model to obtain the updated first feature information of each training sample (i.e. Fig.10 The feature information of the training samples used in generating the second loss function).

[0214] like Fig.10 As shown, the second loss function not only indicates the similarity between the feature information of different positive training samples, but also indicates the similarity between the feature information of different negative training samples, and also indicates the similarity between the feature information of the positive training sample and the feature information of the negative training sample. It should be understood that Fig.10 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0215] In an embodiment of the present application, a second loss function is used to improve the similarity between feature information of different positive training samples generated by the first machine learning model during the feature update process, thereby improving the first machine learning model's ability to understand different positive samples, which is conducive to further improving the accuracy of the prediction information generated by the first machine learning model.

[0216] In addition, using the second loss function to train the first machine learning model can also improve the similarity between the feature information of different negative training samples, and reduce the similarity between the feature information of positive training samples and the feature information of negative training samples, which is beneficial to improving the first machine learning model's ability to understand different negative samples, and is beneficial to improving the first machine learning model's ability to distinguish between positive samples and negative samples, which is beneficial to further improve the accuracy of the prediction information generated by the first machine learning model.

[0217] In order to understand this solution more intuitively, the detailed implementation steps of the training phase of the first machine learning model are explained below with reference to specific examples. Fig.11 In the example, the first information is used as the alarm information, the alarm information carries the alarm type, and the first fault is specifically manifested as a "base station out-of-service fault". The detailed implementation steps of the training phase of the first machine learning model are introduced. The training samples of the first machine learning model are obtained based on the historical alarm data of multiple base stations. Each training sample includes the alarm information generated within M time units. The true value corresponding to each training sample indicates whether a base station out-of-service fault will occur within K time units after M time units. Fig.11 In the embodiment shown, M time units are 7 days and K time units are 3 days. Fig.11 , Fig.11 Another flowchart of the data processing method provided in the embodiment of the present application.

[0218] 1101. Obtain multiple training samples and the true value corresponding to each training sample, each training sample includes at least one alarm type generated within 7 days, and the true value corresponding to each training sample indicates whether a base station out-of-service failure will occur within the next 3 days.

[0219] In the embodiment of the present application, the step of "building training samples based on historical alarm data" can be performed by a training device as an example. The training device obtains the historical alarm data of each of the multiple base stations, and the historical alarm data of each base station is sorted from early to late according to the generation time of the alarm type. For the convenience of description, any one of the multiple base stations is referred to as a "target base station". Optionally, for the historical alarm data collected by the target base station, a training sample can be constructed in a sliding window manner.

[0220] Exemplarily, since the length of M time units is 7 days and the length of K time units is 3 days, the length of the entire sample window is 10 days, and the window is slid according to the specified sliding time interval T (set to 1 day here) to obtain multiple second samples from the historical alarm data collected by the target base station; the training device performs the above operations on the historical alarm data collected by each base station, so that multiple second samples can be obtained based on the historical alarm data collected by multiple base stations.

[0221] For a more intuitive understanding of the "second sample", see Fig.12 , Fig.12 A schematic diagram of a second sample provided in an embodiment of the present application, Fig.12 Take M time units as 7 days and K time units as 3 days as an example. Fig.12 In the example, Alarm A represents alarm information A, Alarm B represents alarm information B, and Alarm C represents alarm information C. Fig.12 As shown, each second sample includes alarm information generated within 10 days, the alarm information generated within 7 days in the second sample is used to determine the training sample, and the alarm information generated within 3 days in the second sample is used to determine the true value corresponding to the training sample. It should be understood that Fig.12 The examples are only for facilitating the understanding of this solution and are not intended to limit this solution.

[0222] After obtaining the plurality of second samples, the training device may label each second sample, that is, label each second sample as a positive sample or a negative sample. The meanings of “positive sample” and “negative sample” can be found in the above Figure 4 The description of step 401 in the corresponding embodiment is not repeated here.

[0223] The training device may determine the N alarm types most relevant to the first fault (i.e., the base station out-of-service fault) using a frequent item mining algorithm based on multiple positive second samples, where N is set to 50 as an example. After determining the 50 alarm types most relevant to the base station out-of-service fault, the training device may obtain multiple training samples based on the multiple second samples.

[0224] Exemplarily, any second sample is referred to as a "target sample" here, and the historical alarm types within 7 days included in the target sample are screened to obtain the alarm types within 7 days included in the target training sample, that is, the target training sample is obtained. Among them, all alarm types in the target training sample are included in the aforementioned 50 alarm types, and other alarm types outside the aforementioned 50 alarm types have been eliminated. The training device performs the aforementioned operation on each second sample, and multiple training samples can be obtained.

[0225] The training device also sets the number of alarm types included in each training sample to 50. If the number of alarm types included in the target training sample is greater than 50, at least one alarm type in the target training sample can be deleted so that the number of alarm types included in the updated target training sample is equal to 50; if the number of alarm types included in the target training sample is less than 50, meaningless alarm types can be added to the target training sample so that the number of alarm types included in the updated target training sample is equal to 50.

[0226] Exemplarily, if the target sample is a positive sample, the true value corresponding to the target training sample indicates that a base station out-of-service failure will occur within the next three days; if the target sample is a negative sample, the true value corresponding to the target training sample indicates that a base station out-of-service failure will not occur within the next three days; it should be noted that the method for obtaining the true value corresponding to each training sample in multiple training samples can refer to the method for obtaining the aforementioned "true value corresponding to the target training sample".

[0227] Exemplarily, if the target sample is a positive sample, the target training sample is a positive training sample, and if the target sample is a negative sample, the target training sample is a negative training sample.

[0228] 1102. Obtain first time information corresponding to a first training sample, where the first training sample is any one of a plurality of training samples, and the first time information includes a first time interval between a time point at which each alarm type is generated and a first time point, where the first time point is an end time point of the first training sample.

[0229] In an embodiment of the present application, the training device can obtain a batch of training samples each time it trains the first machine learning model, and any training sample in a batch of training samples is referred to as a "first training sample". The first training sample includes at least one alarm type generated within a third time period, the time length of the third time period is 7 days, and the first time point is the end time point of the third time period. After obtaining the first training sample, the training device can determine the relative time interval (that is, the first time interval) between each alarm type generated within the third time period and the end time point of the third time period. In this embodiment, the granularity of the relative time interval is taken as hours as an example, that is, the hourly time interval between the generation time of each alarm type in the first training sample and the first time point is determined, so that the first time information corresponding to the first training sample can be obtained.

[0230] The training device performs the above steps on each training sample in a batch of training samples, and can obtain the first time information corresponding to each training sample in a batch of training samples.

[0231] 1103. Obtain first characteristic information of a first training sample, where the first characteristic information includes characteristic information of each alarm type in the first training sample and characteristic information of a first time interval corresponding to each alarm type.

[0232] In the embodiment of the present application, after acquiring the first training sample and the first time information corresponding to the first training sample, the training device may acquire the first feature information of the first training sample.

[0233] Exemplarily, the training device may be deployed with second information, which includes characteristic information of each of the 50 alarm types, and also includes characteristic information of "meaningless alarm type"; since the granularity of the first time interval is hours and the length of M time units is 7 days, the second information also includes characteristic information of 24*7=168 types of time intervals.

[0234] After determining at least one alarm type included in the first training sample and the first time interval corresponding to each alarm type in the first training sample, the training device can obtain the initial feature information of each alarm type in the first training sample and the initial feature information of the first time interval corresponding to each alarm type from the second information to obtain the first feature information.

[0235] Exemplarily, the first feature information includes at least one second feature information corresponding to at least one alarm type in the first training sample; for the convenience of description, any one of the at least one alarm type included in the first training sample is referred to as the "target alarm type", and the training device combines the initial feature information of the target alarm type with the "initial feature information of the first time interval corresponding to the target alarm type" to obtain the second feature information corresponding to the target alarm type in the first feature information. The training device repeats the above steps at least once to obtain the second feature information corresponding to each alarm type in the first training sample, that is, to obtain the first feature information.

[0236] The training device repeatedly executes the above steps multiple times to obtain the first feature information of each training sample in a batch of training samples.

[0237] 1104. Input the first feature information of the first training sample into the first machine learning model to obtain prediction information corresponding to the first training sample output by the first machine learning model, where the prediction information indicates whether a base station decommissioning failure will occur in the next three days.

[0238] In an embodiment of the present application, after the training device obtains the first feature information of each training sample in a batch of training samples, the first feature information of each training sample in a batch of training samples can be input into the first machine learning model in sequence to obtain the prediction information corresponding to each of the aforementioned training samples output by the first machine learning model.

[0239] Exemplarily, after the training device inputs the first feature information of the first training sample into the first machine learning model, the first feature information of the first training sample can be updated through the first machine learning model to obtain the updated first feature information of the first training sample; the training device performs feature processing on the updated first feature information through the first machine learning model to obtain prediction information corresponding to the first training sample output by the first machine learning model.

[0240] Optionally, the first machine learning model may include multiple first feature update modules, and the training device may update the first feature information of the first training sample through the first machine learning model, which may include: the training device may update the first feature information through the first feature update module in the first machine learning model.

[0241] Exemplarily, the training device updates the feature of the first feature information through the first feature update module in the first machine learning model, which may include: the training device divides the multiple second feature information corresponding one-to-one to the multiple alarm types in the first training sample into three first groups, the three first groups include a first group A, a first group B and a first group C, at least one second feature information included in the first group A corresponds to the alarm type generated in the first three days within 7 days, at least one second feature information included in the first group B corresponds to the alarm type generated in the middle two days within 7 days, and at least one second feature information included in the first group C corresponds to the alarm type generated in the last two days within 7 days. The training device updates the features of each second feature information in the first group A based on the self-attention mechanism according to at least one second feature information in the first group A, and obtains the updated feature 1 of each second feature information in the first group A; the training device updates the features of each second feature information in the first group B based on the self-attention mechanism according to at least one second feature information in the first group B, and obtains the updated feature 1 of each second feature information in the first group B; the training device updates the features of each second feature information in the first group C based on the self-attention mechanism according to at least one second feature information in the first group C, and obtains the updated feature 1 of each second feature information in the first group C.

[0242] The training device divides the multiple second feature information corresponding to the multiple alarm types in the first training sample into a second group and a third group, and at least one alarm type corresponding to at least one second feature information included in the second group can directly indicate a base station out-of-service failure; and the third group includes second feature information corresponding to all alarm types other than the second group in the first training sample. The training device updates the features of each second feature information in the second group based on the self-attention mechanism according to at least one second feature information in the second group, and obtains the updated features 2 of each second feature information in the second group; and updates the features of each second feature information in the third group based on the self-attention mechanism according to at least one second feature information in the third group, and obtains the updated features 2 of each second feature information in the third group.

[0243] The training device divides the multiple second feature information corresponding to the multiple alarm types in the first training sample into multiple fourth groups, and the alarm types corresponding to all the second feature information in each fourth group are the same. The training device updates the feature of at least one second feature information in each fourth group based on the self-attention mechanism, and obtains the updated feature 3 of each second feature information in each fourth group.

[0244] The training device combines the updated feature 1, updated feature 2, and updated feature 3 of each second feature information corresponding to the first training sample to obtain the updated second feature information corresponding to each second feature information, that is, the updated first feature information corresponding to the first training sample. It should be noted that the first machine learning model may include one or more first feature update modules. If the first machine learning model includes multiple first feature update modules, the first feature information of the first training sample can be updated multiple times through multiple first feature update modules.

[0245] The training device may perform the above operation on each training sample in a batch of training samples to obtain updated first feature information of each training sample in a batch of training samples.

[0246] 1105. The first machine learning model is trained according to a first loss function and a second loss function, wherein the first loss function indicates the similarity between the predicted information and the true value corresponding to the first training sample, and the goals of training with the second loss function include improving the similarity between the feature information of the positive training samples and reducing the similarity between the feature information of the positive training samples and the feature information of the negative training samples.

[0247] In the embodiments of the present application, for example, both the first loss function and the second loss function may adopt cross entropy (CE), or the first loss function and the second loss function may adopt L1 loss function, L2 loss function or other types of loss functions, etc., or the first loss function and the second loss function may adopt different loss functions, etc., which are not limited in the embodiments of the present application.

[0248] Exemplarily, in this embodiment, a batch of training samples may include 256 training samples. The training device may generate a function value of a first loss function based on the prediction information corresponding to each training sample in a batch of training samples generated in step 1103, and the true value corresponding to each training sample in a batch of training samples. The first loss function indicates the similarity between the prediction information corresponding to each training sample in a batch of training samples and the true value.

[0249] There may be positive training samples and negative training samples in the 256 training samples. The training device may also generate a function value of a second loss function based on the feature information corresponding to each training sample in a batch of training samples generated in step 1103, where the second loss function indicates the similarity between feature information of different positive training samples in a batch of training samples, and the similarity between feature information of positive training samples and feature information of negative training samples in a batch of training samples, wherein “feature information corresponding to each training sample in a batch of training samples” may be understood as the updated first feature information of each training sample in a batch of training samples.

[0250] The training device performs weighted summation on the function value of the first loss function and the function value of the second loss function to obtain a total loss function value; according to the total loss function value, based on the back propagation algorithm, the parameters of the first machine learning model can be gradient updated; optionally, according to the total loss function value, based on the back propagation algorithm, the second information can also be gradient updated.

[0251] To further understand the present solution, an example of a formula used in the process of “weighted summing the function value of the first loss function and the function value of the second loss function to obtain a total loss function value” is disclosed as follows:

[0252] Loss total =α*Loss CE +(1-α)*Loss InfoNCE ;

[0253] Among them, Loss total Represents the total loss function value, LossCE Represents the first loss function, α is a hyperparameter, Loss InfoNCE Represents the second loss function. It should be understood that the example here is only for the convenience of understanding this solution and is not used to limit this solution.

[0254] It should be noted that the training device can repeat steps 1102 to 1105 multiple times to implement iterative updates of the first machine learning model (optionally, also including the second information) until the convergence condition is met to obtain the trained first machine learning model; optionally, the trained second information can also be obtained. Exemplarily, the "convergence condition" can be the convergence condition of satisfying the first loss function and the second loss function, or it can also be the number of training times for the first machine learning model reaching a preset number of times, or other convergence conditions, etc., which are not limited in the embodiments of the present application.

[0255] 2. Reasoning Stage

[0256] For details, please refer to Fig.13 , Fig.13 Another schematic diagram of the data processing method provided in the embodiment of the present application is as follows: Fig.13 As shown, the data processing method provided by the present application may include:

[0257] 1301. Obtain a first sample, where the first sample includes at least one first information generated in a first time period.

[0258] 1302. Obtain first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point.

[0259] 1303. Input the first feature information into the first machine learning model, update the features of the first feature information through the first machine learning model, and obtain prediction information corresponding to the second time period output by the first machine learning model, wherein the first feature information includes the feature information of the first information and the feature information of the first time interval corresponding to the first information, and the second time period is located after the first time period.

[0260] In the embodiment of the present application, Fig.13 The first machine learning model in the corresponding embodiment is Figures 4 to 12 The meaning of "first sample" is similar to that of "first training sample", except that the third time period in the first training sample is replaced by the first time period in the first sample; the meaning of "prediction information corresponding to the second time period" is similar to that of "prediction information corresponding to the fourth time period", except that the "fourth time period" is replaced by the "second time period". Fig.13 The meanings of the terms in the corresponding embodiments can be found in the above Figures 3 to 13 The specific implementation of steps 1301 to 1303 can also refer to the above Figures 3 to 13 The description in the corresponding embodiment is not repeated here.

[0261] In order to have a more intuitive understanding of the beneficial effects of the method provided by the present application, the beneficial effects of the method provided by the present application are described below in combination with experimental data. In this experiment, the first information includes at least one alarm type collected by the base station within the first time period, and the prediction information indicates whether a certain fault will occur in the base station within the second time period. For the experimental results, please refer to the following table.

[0262]

[0263] Table 1

[0264] As shown in Table 1 above, the performance evaluation indicators used include Precision: recall = 0.3 and Precision: recall = 0.2, where Precision represents the ratio of the number of samples that can correctly predict whether a fault will occur to the total number of samples predicted to fail, and recall represents the ratio of the number of samples that can correctly predict whether a fault will occur to the total number of samples in the test set that will fail. During the experiment, the historical alarm data of base station equipment supplier 1 and the historical alarm data of supplier 2 were used respectively to predict whether the base station will send a certain fault in the second time period. It can be seen from Table 1 above that the method provided in the present application can most accurately predict whether a certain fault will occur in the base station in the future second time period.

[0265] exist Figures 1a to 13 On the basis of the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Fig.14 , Fig.14 A structural schematic diagram of a data processing device provided for an embodiment of the present application, the data processing device 1400 includes: an acquisition module 1401, used to acquire at least one first information generated within a first time period; the acquisition module 1401 is also used to acquire first time information, the first time information includes a first time interval between a time point when the first information is generated and the first time point, and the first time point is a preset reference time point; an input module 1402, used to input the first feature information into a machine learning model to obtain prediction information corresponding to a second time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of a first time interval corresponding to the first information, and the second time period is located after the first time period.

[0266] Optionally, the first time point is determined based on a start time point or an end time point of the first time period.

[0267] Optionally, the first information includes alarm information of the network device, the prediction information corresponding to the second time period indicates whether a first fault occurs in the network device within the second time period, and the time point when the first information is generated is the time point when the alarm information occurs.

[0268] Optionally, the first information includes physiological information of the user, the prediction information corresponding to the second time period indicates the motion state of the user in the second time period, and the time point when the first information is generated is the time point when the physiological information occurs.

[0269] Optionally, the first feature information includes at least one second feature information corresponding to at least one first information, and each second feature information includes feature information of one first information and feature information of a first time interval corresponding to one first information. In the process of updating the features of the first feature information through the machine learning model, based on each of the at least two grouping criteria, at least one second feature information is divided into at least two groups, and the data processing device further includes: an updating module 1403, which is used to update the second feature information in the same group according to the second feature information in the same group of the at least two groups.

[0270] Optionally, the first information includes an alarm type, and the prediction information corresponding to the second time period indicates whether a first fault occurs within the second time period; the at least two grouping criteria include any at least two of the following: the generation time of the first information, whether the alarm type included in the first information can directly indicate the first fault, or the alarm type in the first information.

[0271] Optionally, the first information is warning information, and the prediction information corresponding to the second time period indicates whether the first fault occurs in the second time period. The step of updating the second characteristic information in the same group according to the second characteristic information in the same group of at least two groups includes any of the following:

[0272] updating the second characteristic information in the same first group according to the second characteristic information in the same first group in the at least two first groups, wherein when a first grouping basis in the at least two grouping basis is adopted, the at least two groups are at least two first groups, based on the first grouping basis, the first time period is divided into at least two sub-time periods, and each first group includes the second characteristic information corresponding to all the first information generated in one sub-time period in the at least two sub-time periods; or,

[0273] The second characteristic information in the second group is updated according to the second characteristic information in the second group, and the second characteristic information in the third group is updated according to the second characteristic information in the third group, wherein when the second grouping basis among the at least two grouping basis is adopted, the at least two groups include the second group and the third group, the second group includes the second characteristic information corresponding to all the first information in the at least one first information that can directly indicate the first fault, and the third group includes the second characteristic information corresponding to all the first information in the at least one first information except the second group; or,

[0274] According to the second characteristic information in the same fourth group in at least two fourth groups, the second characteristic information in the same fourth group is updated, wherein, when the third grouping basis among at least two grouping basis is adopted, at least two groups are at least two fourth groups, and the alarm type of the first information corresponding to all the second characteristic information in each fourth group is the same.

[0275] Optionally, the acquisition module 1401 is also used to acquire characteristic information of the first information and characteristic information of the first time interval corresponding to the first information based on pre-stored second information, so as to obtain the first characteristic information, wherein the second information indicates the characteristic information corresponding to each of the multiple alarm information, and the second information also indicates the characteristic information corresponding to each of the multiple time intervals, and the second information is updated during the process of executing the training operation of the machine learning model; the characteristic information corresponding to the alarm information is obtained based on the feature extraction operation performed on the alarm information, and the characteristic information corresponding to the time interval is obtained based on the feature extraction operation performed on the time interval.

[0276] Optionally, the training process of the machine learning model adopts multiple training samples, the true value corresponding to each training sample and the loss function, the training sample includes at least one first information generated in a third time period, the true value corresponding to the training sample indicates whether a first fault occurred in a fourth time period, and the fourth time period is located after the third time period; wherein, when the true value corresponding to the training sample indicates that a first fault occurred in the fourth time period, the training sample is a positive training sample, the loss function is determined based on the feature information of the positive training sample, and the goal of training using the loss function includes improving the similarity between the feature information of different positive training samples.

[0277] Optionally, when the true value corresponding to the training sample indicates that the first fault did not occur within the fourth time period, the training sample is a negative training sample, and the loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample. The goal of training using the loss function also includes: improving the similarity between the feature information of different negative training samples, and / or reducing the similarity between the feature information of the positive training sample and the feature information of the negative training sample.

[0278] It should be noted that the information interaction, execution process, etc. between the modules / units in the data processing device 1400 are the same as those in the present application. Figures 3 to 13 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0279] See also Fig.15 , Fig.15 Another structural schematic diagram of a data processing device provided for an embodiment of the present application, the data processing device 1500 includes: an acquisition module 1501, used to acquire training samples, the training samples include at least one first information generated within a third time period; the acquisition module 1501 is also used to acquire first time information, the first time information includes a first time interval between a time point when the first information is generated and the first time point, and the first time point is a preset reference time point; an input module 1502, used to input the first feature information into a machine learning model to obtain prediction information corresponding to a fourth time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of the first time interval corresponding to the first information, and the fourth time period is located after the third time period; a training module 1503, used to train the machine learning model according to a first loss function, the first loss function indicating the similarity between the prediction information corresponding to the fourth time period and the true value corresponding to the training sample.

[0280] Optionally, the first feature information includes at least one second feature information corresponding one-to-one to at least one first information, and each second feature information includes feature information of a first information and feature information of a first time interval corresponding to a first information; wherein, in the process of updating the features of the first feature information through the machine learning model, based on each of at least two grouping criteria, at least one second feature information is divided into at least two groups, and the data processing device also includes: an updating module, which is used to update the feature information of the first information in the same group according to the feature information of the first information in the same group of at least two groups.

[0281] Optionally, the first information is alarm information, and the prediction information corresponding to the fourth time period indicates whether a first fault occurs within the fourth time period; the acquisition module 1501 is also used to obtain characteristic information of the first information and characteristic information of a first time interval corresponding to the first information based on the second information, wherein the second information indicates characteristic information corresponding to each type of alarm information in a plurality of alarm information, and the second information also indicates characteristic information corresponding to each time interval in a plurality of time intervals, the characteristic information corresponding to the alarm information is obtained based on a feature extraction operation performed on the alarm information, and the characteristic information corresponding to the time interval is obtained based on a feature extraction operation performed on the time interval; the training module 1503 is specifically used to update the weight parameters and the second information of the machine learning model according to the first loss function.

[0282] Optionally, when the true value corresponding to the training sample indicates that a first fault occurred within a fourth time period, the training sample is a positive training sample, wherein the training module 1503 is specifically used to train the machine learning model according to the first loss function and the second loss function, the second loss function is determined based on the feature information of the positive training sample, and the goal of training using the second loss function includes improving the similarity between the feature information of different positive training samples.

[0283] Optionally, when the true value corresponding to the training sample indicates that the first fault did not occur within a fourth time period, the training sample is a negative training sample, wherein the second loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample, and the goal of training using the second loss function also includes: improving the similarity between the feature information of different negative training samples, and / or reducing the similarity between the feature information of the positive training sample and the feature information of the negative training sample.

[0284] It should be noted that the information interaction, execution process, etc. between the modules / units in the data processing device 1500 are the same as those in the present application. Figures 3 to 13 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0285] See also Fig.16 , Fig.16Another structural schematic diagram of a data processing device provided for an embodiment of the present application, the data processing device 1600 includes: an acquisition module 1601, used to acquire at least one alarm information generated within a first time period; an input module 1602, used to input first feature information into a machine learning model, the first feature information including at least one second feature information corresponding one-to-one to the at least one alarm information; an update module 1603, used to update the second feature information in the same group according to the second feature information in the same group of at least two groups, and obtain prediction information corresponding to the second time period output by the machine learning model, wherein, in the process of updating the second feature information through the machine learning model, based on each of the at least two grouping criteria, the at least one second feature information is divided into at least two groups, and the second time period is located after the first time period.

[0286] Optionally, the acquisition module 1601 is also used to obtain first time information, the first time information includes a first time interval between the time point when the alarm information is generated and the first time point, the first time point is a preset reference time point, and each second characteristic information includes characteristic information of an alarm information and characteristic information of the first time interval corresponding to an alarm information.

[0287] It should be noted that the information interaction, execution process, etc. between the modules / units in the data processing device 1600 are the same as those in the present application. Figures 3 to 13 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0288] Next, an execution device provided by an embodiment of the present application is introduced. Fig.17 , Fig.17 A schematic diagram of a structure of an execution device provided in an embodiment of the present application, specifically, the execution device 1700 includes: a receiver 1701, a transmitter 1702, a processor 1703 and a memory 1704 (wherein the number of processors 1703 in the execution device 1700 can be one or more, Fig.17 In the example of FIG. 1701 , a processor 1703 is used, where the processor 1703 may include an application processor 17031 and a communication processor 17032. In some embodiments of the present application, the receiver 1701, the transmitter 1702, the processor 1703 and the memory 1704 may be connected via a bus or other means.

[0289] The memory 1704 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1703. A portion of the memory 1704 may also include a non-volatile random access memory (NVRAM). The memory 1704 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0290] The processor 1703 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.

[0291] The method disclosed in the above embodiment of the present application can be applied to the processor 1703, or implemented by the processor 1703. The processor 1703 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1703. The above processor 1703 can be a general processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The processor 1703 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be combined and executed. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1704, and the processor 1703 reads the information in the memory 1704 and completes the steps of the above method in combination with its hardware.

[0292] The receiver 1701 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the execution device. The transmitter 1702 can be used to output digital or character information through the first interface; the transmitter 1702 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1702 can also include a display device such as a display screen.

[0293] In the embodiment of the present application, the processor 1703 is used to execute Figure 3 or Fig.13 The data processing method executed by the execution device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 17031 in the processor 1703 executes the above steps is the same as that in the present application. Figures 3 to 13 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 13 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0294] The present application also provides a training device. Fig.18 , Fig.18 It is a structural diagram of a training device provided in an embodiment of the present application. Specifically, the training device 1800 is implemented by one or more servers. The training device 1800 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1822 (for example, one or more processors) and memory 1832, and one or more storage media 1830 (for example, one or more mass storage devices) storing application programs 1842 or data 1844. Among them, the memory 1832 and the storage medium 1830 can be short-term storage or permanent storage. The program stored in the storage medium 1830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the training device. Furthermore, the central processing unit 1822 can be configured to communicate with the storage medium 1830 to execute a series of instruction operations in the storage medium 1830 on the training device 1800.

[0295] The training device 1800 may also include one or more power supplies 1826, one or more wired or wireless network interfaces 1850, one or more input and output interfaces 1858, and / or, one or more operating systems 1841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0296] In the embodiment of the present application, the central processor 1822 is used to execute Figures 4 to 12 The training method of the model executed by the training device in the corresponding embodiment. It should be noted that the specific manner in which the central processor 1822 executes the above steps is the same as that in the present application. Figures 3 to 13 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 3 to 13 The corresponding method embodiments are the same. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0297] The present application also provides a computer-readable storage medium in which a program for signal processing is stored. When the program is run on a computer, the computer executes the above-mentioned Figures 3 to 13 The method described in the embodiment shown in the figure executes the steps executed by the device, or causes the computer to execute the above-mentioned Figures 3 to 13 The illustrated embodiment describes the steps performed by the training device in the method.

[0298] The present application also provides a computer program product which, when executed on a computer, enables the computer to execute the above Figures 3 to 13 The method described in the embodiment shown in the figure executes the steps executed by the device, or causes the computer to execute the above-mentioned Figures 3 to 13 The illustrated embodiment describes the steps performed by the training device in the method.

[0299] The execution device or training device provided in the embodiment of the present application may be a chip, which includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit. The processing unit may execute the computer execution instructions stored in the storage unit to enable the chip to execute the above Figures 3 to 13 The data processing method described in the embodiment shown. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0300] For details, please refer to Fig.19 , Fig.19A schematic diagram of the structure of a chip provided in an embodiment of the present application, the chip can be expressed as a neural network processor NPU 190, NPU 190 is mounted on the host CPU (Host CPU) as a coprocessor, and the Host CPU assigns tasks. The core part of the NPU is the operation circuit 1903, which is controlled by the controller 1904 to extract matrix data from the memory and perform multiplication operations.

[0301] In some implementations, the operation circuit 1903 includes multiple processing units (Process Engine, PE) inside. In some implementations, the operation circuit 1903 is a two-dimensional systolic array. The operation circuit 1903 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1903 is a general-purpose matrix processor.

[0302] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 1902 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1901 and performs matrix operations with matrix B. The partial results or final results of the obtained matrix are stored in the accumulator 1908.

[0303] The unified memory 1906 is used to store input data and output data. The weight data is directly transferred to the weight memory 1902 through the direct memory access controller (DMAC) 1905. The input data is also transferred to the unified memory 1906 through the DMAC.

[0304] BIU stands for Bus Interface Unit 1919 , which is used for the interaction between the AXI bus and the DMAC and instruction fetch buffer (IFB) 1909 .

[0305] The bus interface unit 1919 (Bus Interface Unit, BIU for short) is used for the instruction fetch memory 1909 to obtain instructions from the external memory, and is also used for the storage unit access controller 1905 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0306] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1906 or to transfer weight data to the weight memory 1902 or to transfer input data to the input memory 1901.

[0307] The vector calculation unit 1907 includes multiple operation processing units, and further processes the output of the operation circuit when necessary, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of feature planes, etc.

[0308] In some implementations, the vector calculation unit 1907 can store the processed output vector to the unified memory 1906. For example, the vector calculation unit 1907 can apply a linear function and / or a nonlinear function to the output of the operation circuit 1903, such as linear interpolation of the feature plane extracted by the convolution layer, and then, for example, a vector of accumulated values ​​to generate an activation value. In some implementations, the vector calculation unit 1907 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1903, for example, for use in a subsequent layer in a neural network.

[0309] An instruction fetch buffer 1909 connected to the controller 1904 is used to store instructions used by the controller 1904;

[0310] Unified memory 1906, input memory 1901, weight memory 1902 and instruction fetch memory 1909 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0311] Among them, the operations of each layer in the machine learning model in the above embodiment can be performed by the operation circuit 1903 or the vector calculation unit 1907.

[0312] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect method.

[0313] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0314] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0315] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0316] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire at least one first information generated within a first time period; Acquire first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point; The first feature information is input into a machine learning model to obtain prediction information corresponding to a second time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of a first time interval corresponding to the first information, and the second time period is located after the first time period.

2. The method according to claim 1, characterized in that The first time point is determined based on a start time point or an end time point of the first time period.

3. The method according to claim 1 or 2, characterized in that: The first information includes alarm information of the network device, the prediction information corresponding to the second time period indicates whether a first fault occurs in the network device within the second time period, and the time point when the first information is generated is the time point when the alarm information occurs.

4. The method according to claim 1 or 2, characterized in that: The first information includes physiological information of the user, the prediction information corresponding to the second time period indicates the motion state of the user in the second time period, and the time point of generating the first information is the time point of occurrence of the physiological information.

5. The method according to claim 1 or 2, characterized in that: The first characteristic information includes at least one second characteristic information corresponding to the at least one first information, and each second characteristic information includes characteristic information of one first information and characteristic information of a first time interval corresponding to the one first information; Wherein, in the process of updating the first feature information by using the machine learning model, based on each of at least two grouping criteria, the at least one second feature information is divided into at least two groups, and the method further includes: According to the second characteristic information in the same group of the at least two groups, the second characteristic information in the same group is updated.

6. The method according to claim 5, characterized in that The first information is warning information, and the prediction information corresponding to the second time period indicates whether the first fault occurs in the second time period; The at least two grouping criteria include any at least two of the following: generation time of the first information, whether the alarm type included in the first information can directly indicate the first fault, or the alarm type in the first information.

7. The method according to claim 5, characterized in that The first information is warning information, and the prediction information corresponding to the second time period indicates whether the first fault occurs in the second time period; Wherein, updating the second characteristic information in the same group according to the second characteristic information in the same group of the at least two groups includes any of the following: updating the second characteristic information in the same first group of at least two first groups according to the second characteristic information in the same first group, wherein when a first grouping basis among the at least two grouping basis is adopted, the at least two groups are the at least two first groups, and based on the first grouping basis, the first time period is divided into at least two sub-time periods, and each of the first groups includes the second characteristic information corresponding to all the first information generated in one sub-time period of the at least two sub-time periods; or According to the second characteristic information in the second group, the second characteristic information in the second group is updated, and according to the second characteristic information in the third group, the second characteristic information in the third group is updated, wherein, when the second grouping basis among the at least two grouping basis is adopted, the at least two groups include the second group and the third group, the second group includes the second characteristic information corresponding to all the first information in the at least one first information that can directly indicate the first fault, and the third group includes the second characteristic information corresponding to all the first information in the at least one first information other than the second group; or, According to the second characteristic information in the same fourth group among at least two fourth groups, the second characteristic information in the same fourth group is updated, wherein when the third grouping basis among the at least two grouping basis is adopted, the at least two groups are the at least two fourth groups, and the alarm type of the first information corresponding to all the second characteristic information in each of the fourth groups is the same.

8. The method according to claim 3, characterized in that Before inputting the first feature information into the machine learning model to obtain prediction information corresponding to the second time period output by the machine learning model, the method further includes: Acquire characteristic information of the first information and characteristic information of the first time interval corresponding to the first information according to the pre-stored second information, so as to obtain the first characteristic information, wherein the second information indicates characteristic information corresponding to each type of alarm information in a plurality of alarm information, and the second information further indicates characteristic information corresponding to each time interval in a plurality of time intervals, and the second information is updated during the process of executing the training operation of the machine learning model; The feature information corresponding to the alarm information is obtained based on a feature extraction operation performed on the alarm information, and the feature information corresponding to the time interval is obtained based on a feature extraction operation performed on the time interval.

9. The method according to claim 1 or 2, characterized in that: The training process of the machine learning model uses a plurality of training samples, a true value corresponding to each of the training samples, and a loss function, the training samples include at least one first information generated in a third time period, the true value corresponding to the training sample indicates whether a first fault occurs in a fourth time period, and the fourth time period is located after the third time period; Among them, when the true value corresponding to the training sample indicates that the first fault occurred within the fourth time period, the training sample is a positive training sample, the loss function is determined based on the feature information of the positive training sample, and the goal of training using the loss function includes improving the similarity between the feature information of different positive training samples.

10. The method according to claim 9, characterized in that When the true value corresponding to the training sample indicates that the first fault did not occur within the fourth time period, the training sample is a negative training sample, and the loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample. The goal of training using the loss function also includes: improving the similarity between feature information of different negative training samples, and / or reducing the similarity between feature information of positive training samples and feature information of negative training samples.

11. A data processing method, characterized in that: The method comprises: Acquire a training sample, where the training sample includes at least one first information generated within a third time period; Acquire first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point; Inputting the first feature information into a machine learning model to obtain prediction information corresponding to a fourth time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of the first time interval corresponding to the first information, and the fourth time period is after the third time period; The machine learning model is trained according to a first loss function, wherein the first loss function indicates a similarity between the prediction information corresponding to the fourth time period and a true value corresponding to the training sample.

12. The method according to claim 11, characterized in that The first characteristic information includes at least one second characteristic information corresponding to the at least one first information, and each second characteristic information includes characteristic information of one first information and characteristic information of a first time interval corresponding to the one first information; Wherein, in the process of updating the first feature information by using the machine learning model, based on each of at least two grouping criteria, the at least one second feature information is divided into at least two groups, and the method further includes: According to the characteristic information of the first information in the same group of the at least two groups, the characteristic information of the first information in the same group is updated.

13. The method according to claim 11 or 12, characterized in that: The first information is warning information, the prediction information corresponding to the fourth time period indicates whether a first fault occurs in the fourth time period, and before inputting the first feature information into the machine learning model, the method further includes: Acquire, according to the second information, characteristic information of the first information and characteristic information of the first time interval corresponding to the first information, wherein the second information indicates characteristic information corresponding to each type of alarm information among multiple types of alarm information, and the second information further indicates characteristic information corresponding to each time interval among multiple time intervals, the characteristic information corresponding to the alarm information is obtained based on a feature extraction operation performed on the alarm information, and the characteristic information corresponding to the time interval is obtained based on a feature extraction operation performed on the time interval; The training of the machine learning model according to the first loss function includes: updating the weight parameters of the machine learning model and the second information according to the first loss function.

14. The method according to claim 11 or 12, characterized in that: When the true value corresponding to the training sample indicates that a first fault has occurred within the fourth time period, the training sample is a positive training sample, wherein the training of the machine learning model according to the first loss function includes: The machine learning model is trained according to the first loss function and the second loss function, wherein the second loss function is determined according to the feature information of the positive training sample, and the goal of training using the second loss function includes improving the similarity between the feature information of different positive training samples.

15. The method according to claim 14, characterized in that When the true value corresponding to the training sample indicates that the first fault did not occur within the fourth time period, the training sample is a negative training sample, wherein the second loss function is determined based on the feature information of the positive training sample and the feature information of the negative training sample, and the goal of training using the second loss function also includes: improving the similarity between feature information of different negative training samples, and / or reducing the similarity between feature information of positive training samples and feature information of negative training samples.

16. A data processing method, characterized in that: The method comprises: Obtain at least one alarm information generated within a first time period; Inputting first feature information into a machine learning model, wherein the first feature information includes at least one second feature information corresponding one-to-one to the at least one warning information; Based on the second feature information in the same group of at least two groups, the second feature information in the same group is updated to obtain prediction information corresponding to the second time period output by the machine learning model, wherein, in the process of updating the second feature information through the machine learning model, based on each of at least two grouping criteria, the at least one second feature information is divided into the at least two groups, and the second time period is located after the first time period.

17. The method according to claim 16, characterized in that The method further comprises: Obtain first time information, the first time information including a first time interval between a time point at which the alarm information is generated and a first time point, the first time point being a preset reference time point, and each of the second characteristic information including characteristic information of one alarm information and characteristic information of the first time interval corresponding to the one alarm information.

18. A data processing device, characterized in that: The device comprises: An acquisition module, used to acquire at least one first information generated within a first time period; The acquisition module is further used to acquire first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point; An input module is used to input first feature information into a machine learning model to obtain prediction information corresponding to a second time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of a first time interval corresponding to the first information, and the second time period is located after the first time period.

19. A data processing device, characterized in that: The device comprises: An acquisition module, configured to acquire a training sample, wherein the training sample includes at least one first information generated within a third time period; The acquisition module is further used to acquire first time information, where the first time information includes a first time interval between a time point when the first information is generated and a first time point, where the first time point is a preset reference time point; an input module, configured to input first feature information into a machine learning model to obtain prediction information corresponding to a fourth time period output by the machine learning model, wherein the first feature information includes feature information of the first information and feature information of the first time interval corresponding to the first information, and the fourth time period is located after the third time period; A training module is used to train the machine learning model according to a first loss function, wherein the first loss function indicates the similarity between the prediction information corresponding to the fourth time period and the true value corresponding to the training sample.

20. A data processing device, characterized in that: The device comprises: An acquisition module, used to acquire at least one alarm information generated within a first time period; An input module, configured to input first feature information into a machine learning model, wherein the first feature information includes at least one second feature information corresponding one-to-one to the at least one warning information; An updating module is used to update the second feature information in the same group of at least two groups based on the second feature information in the same group, and obtain prediction information corresponding to the second time period output by the machine learning model, wherein, in the process of updating the second feature information through the machine learning model, based on each of at least two grouping criteria, the at least one second feature information is divided into the at least two groups, and the second time period is located after the first time period.

21. A device, characterized in that comprising a processor and a memory, the processor being coupled to the memory, The memory is used to store programs; The processor is configured to execute the program in the memory so that the device performs the method according to any one of claims 1 to 17.

22. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, and when the program is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 17.

23. A computer program product, characterized in that The computer program product comprises a program, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 17 .