Method for processing data and apparatus thereof
By layering the time series data and fusing the prediction results, the problem of degradation of prediction accuracy caused by the missing input data is solved, and a more efficient prediction effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/137279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-03
AI Technical Summary
In the existing time series data prediction task, the model input feature dimensions are fixed, which leads to the algorithm failure or the prediction accuracy is low when the input data is missing.
The input data is layered, the corresponding model is trained for each layer of data, and the final prediction is made by fusing the prediction results of multiple models, ensuring that even if some layers of data are missing, prediction can be effectively predicted and the prediction accuracy is improved.
Even if some layers of data are missing, predictions can still be made through unmissed data and corresponding models to ensure the prediction effect and improve the prediction accuracy.
Smart Images

Figure CN2024137279_03072025_PF_FP_ABST
Abstract
Description
A data processing method and device thereof
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 29, 2023, with application number 202311868156.7 and application name “A data processing method and device thereof”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence, and in particular to a data processing method and device thereof. Background Art
[0003] Time series data prediction tasks involve predicting future data based on historical time series data. To achieve more accurate predictions, the input data often consists of multiple categories related to the data being predicted. For example, in coal mining, gas concentration prediction uses data such as gas concentration, wind speed, wind pressure, carbon monoxide, temperature, coal cutting speed, gas extraction volume, or gas extraction pressure as input to predict future gas concentrations.
[0004] Existing prediction tasks are often completed using only a single algorithm model. In this approach, the dimensionality of the model input features is fixed, that is, the number of categories of input data remains fixed during model training and inference. When input data is missing, the corresponding dimension has no input, which will cause the algorithm to fail or the prediction accuracy to be low. Summary of the Invention
[0005] In a first aspect, the present application provides a data processing method, which includes: obtaining first data; sending a first model trained based on the first data to an end side; the first model is a base model for time series prediction; obtaining second data; the second data and the first data are data of different types; sending a second model trained based on the second data to an end side; the second model is used to predict data related to the prediction deviation of the base model.
[0006] The idea of this application is to stratify the input data, wherein each layer of data may include one or more categories of time series data, and a corresponding model may be trained for each layer of data. The final prediction result is based on the fusion of the prediction results of multiple models, so that even if the data of some layers is missing, the non-missing data and the corresponding model can still be used for prediction, thereby ensuring the prediction effect. In addition, since a large number of categories of data are not used as input for the same model, the algorithm model can still be feasible and effective even if some categories of data are missing. In addition, since the value of the data related to the prediction deviation is very small compared to the main body of the predicted data, when the input data of the corresponding category of the model used to predict the data related to the prediction deviation is missing, the impact on the final predicted data is very small, so that when the model is trained or inferred, the model used to predict the main body of the data to be predicted is less likely to be missing, thereby improving the prediction accuracy.
[0007] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0008] In a possible implementation, the first data and the second data include time series data within the same time period. For example, the first data and the second data are time series data within the same time period, or the first data and the second data are time series data within a partially overlapping time period.
[0009] In one possible implementation, the first data and the second data include data collected for the same scene within a first time period, and the first data also includes data collected at a target moment after the first time period; before sending the second model trained based on the second data to the end side, the method further includes: obtaining predicted data for the target moment through the first model based on the data collected within the first time period in the first data; using the residual between the predicted data and the data collected at the target moment in the first data as the first true value, and training the second model based on the data collected within the first time period in the second data.
[0010] In one possible implementation, the true value used in the training of the model corresponding to one or more layers of data can be set to the data to be predicted (for the convenience of description, the model can be called the model for predicting the main body of the data to be predicted, that is, the base model), and the true value used in the training of the model corresponding to the data of other layers can be set to the prediction residual of the above model. The prediction residual can be understood as the prediction result of the difference between the prediction result of the model and the true result.
[0011] In one possible implementation, the method further includes: obtaining third data; the third data, the second data, and the first data are different types of time series data collected for the same scene; sending a third model trained based on the third data to the end side; the third model is used to predict data related to the prediction deviation of the second model.
[0012] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0013] In one possible implementation, the first data, the second data, and the third data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period; before sending the third model trained based on the third data to the end side, the method further includes: obtaining a prediction residual of the first model through the second model based on the data collected within the first time period in the second data; using the residual between the prediction residual and the first true value as the second true value, and training a third model based on the data collected within the first time period in the third data.
[0014] In one possible implementation, the first data is sensor data of a first category, and the second data is sensor data of a second category; wherein the probability of missing data when collecting the sensor data of the first category is lower than the probability of missing data when collecting the sensor data of the second category; or, the first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category. The first category and the third category may be the same or different.
[0015] In a possible implementation, the method further includes: receiving indication information sent by the terminal side; the indication information is used to indicate that the first data is used as a training sample of a base model for the time series prediction.
[0016] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0017] In a second aspect, the present application provides a data processing method, the method comprising:
[0018] sending first data;
[0019] Receive a first model; the first model is trained based on the first data, and the first model is a base model for time series prediction;
[0020] Sending second data; the second data and the first data are different types of data;
[0021] Receive a second model; the second model is trained based on the second data, and the second model is used to predict data related to the prediction deviation of the base model.
[0022] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0023] In a possible implementation, the first data and the second data include time series data within the same time period.
[0024] In one possible implementation, the method further includes:
[0025] Sending third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0026] Receive a third model; the third model is trained based on the third data, and the third model is used to predict data related to the prediction deviation of the second model.
[0027] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0028] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0029] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0030] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0031] In one possible implementation, the method further includes:
[0032] Sending indication information; the indication information is used to indicate that the first data is used as a training sample of the base model of the time series prediction.
[0033] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0034] In a third aspect, the present application provides a data processing method, the method comprising:
[0035] Acquire first data and second data; the first data and the second data are different types of data collected for the same scene within a first time period;
[0036] Target data is obtained based on the first data and the second data; wherein the target data is a fusion result of the first prediction data and the second prediction data, the first prediction data is obtained by predicting the data at the target time after the first time period based on the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data.
[0037] In a possible implementation, the second prediction data is a prediction residual of the first prediction data determined based on the second data.
[0038] In a possible implementation, the first data and the second data are time series data.
[0039] In a possible implementation, obtaining target data according to the first data and the second data includes:
[0040] Obtaining, based on the first data, predicted data for a target time after the first time period;
[0041] obtaining second prediction data according to the second data;
[0042] Target data is determined according to a fusion result of the first prediction data and the second prediction data.
[0043] In a possible implementation, obtaining target data according to the first data and the second data includes:
[0044] The first data and the second data are sent to a server, and target data sent from the server is received.
[0045] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0046] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0047] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0048] In one possible implementation, the method further includes:
[0049] Acquire third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0050] The obtaining target data according to the first data and the second data includes:
[0051] Target data is obtained based on the first data, the second data and the third data; the target data is a fusion result of the first prediction data, the second prediction data and the third prediction data, and the third prediction data is data related to the prediction deviation of the second prediction data determined based on the third data, or the third prediction data is data related to the prediction deviation of the first prediction data determined based on the third data.
[0052] In a possible implementation, the third prediction data is a prediction residual of the first prediction data or the second prediction data determined according to the third data.
[0053] In one possible implementation, the first prediction data is obtained by predicting data at a target time after the first time period based on the first data through a first model, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data through a second model.
[0054] In a possible implementation, before obtaining the first data and the second data, the method further includes:
[0055] Obtain or send fourth data and fifth data to the server; the fourth data and the first data are of the same type, and the fifth data and the second data are of the same type; the first model is trained based on the fourth data, and the second model is trained based on the fifth data.
[0056] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0057] The present application also provides a data processing method, which includes:
[0058] Target data is obtained based on first data and second data; wherein the target data is a fusion result of first prediction data and second prediction data, the first prediction data is obtained by predicting the target class based on the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data.
[0059] In a possible implementation, the priority of the first data is higher than that of the second data, and the priority is related to at least one of the following: relevance to the data of the target class and data quality.
[0060] In a possible implementation, the first model is a base model, and the first data is a basic prediction value.
[0061] In one possible implementation, the method further includes:
[0062] An early warning is issued based on the relationship between the target data and the early warning threshold.
[0063] In a possible implementation, the target data is gas information, and the first data and the second data are data related to the gas information.
[0064] In a possible implementation, the first data and the second data are different types of data collected for the same scene within a first time period; the first predicted data is obtained by predicting data at a target time after the first time period based on the first data.
[0065] In a possible implementation, the second prediction data is a prediction residual of the first prediction data determined based on the second data.
[0066] In a possible implementation, the first data and the second data are time series data.
[0067] In a possible implementation, obtaining target data according to the first data and the second data includes:
[0068] Obtaining, based on the first data, predicted data for a target time after the first time period;
[0069] obtaining second prediction data according to the second data;
[0070] Target data is determined according to a fusion result of the first prediction data and the second prediction data.
[0071] In a possible implementation, obtaining target data according to the first data and the second data includes:
[0072] The first data and the second data are sent to a server, and target data sent from the server is received.
[0073] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0074] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0075] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0076] In one possible implementation, the method further includes:
[0077] Acquire third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0078] The obtaining target data according to the first data and the second data includes:
[0079] Target data is obtained based on the first data, the second data and the third data; the target data is a fusion result of the first prediction data, the second prediction data and the third prediction data, and the third prediction data is data related to the prediction deviation of the second prediction data determined based on the third data, or the third prediction data is data related to the prediction deviation of the first prediction data determined based on the third data.
[0080] In a possible implementation, the third prediction data is a prediction residual of the first prediction data or the second prediction data determined according to the third data.
[0081] In one possible implementation, the first prediction data is obtained by predicting data at a target time after the first time period based on the first data through a first model, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data through a second model.
[0082] In a possible implementation, before obtaining the first data and the second data, the method further includes:
[0083] Obtain or send fourth data and fifth data to the server; the fourth data and the first data are of the same type, and the fifth data and the second data are of the same type; the first model is trained based on the fourth data, and the second model is trained based on the fifth data.
[0084] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0085] In a fourth aspect, the present application provides a data processing device, comprising:
[0086] A processing module, configured to obtain first data; and obtain second data; wherein the second data and the first data are different types of data;
[0087] The transceiver module is used to send a second model trained based on the second data to the end side; the second model is used to predict data related to the prediction deviation of the base model.
[0088] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0089] In a possible implementation, the first data and the second data include time series data within the same time period.
[0090] In a possible implementation, the first data and the second data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period;
[0091] Before sending the second model trained according to the second data to the end side, the processing module is further configured to:
[0092] Obtaining predicted data for the target time using the first model based on the data collected during the first time period in the first data;
[0093] The residual between the predicted data and the data collected at the target time in the first data is used as a first true value, and a second model is trained based on the data collected during the first time period in the second data.
[0094] In a possible implementation, the processing module is further configured to:
[0095] Acquire third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0096] The transceiver module is further used to: send a third model trained based on the third data to the end side; the third model is used to predict data related to the prediction deviation of the second model.
[0097] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0098] In a possible implementation, the first data, the second data, and the third data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period;
[0099] The processing module is further configured to obtain, by using the second model, a prediction residual of the first model based on data collected during the first time period in the second data, before sending the third model trained based on the third data to the end side;
[0100] The residual between the prediction residual and the first true value is used as the second true value, and a third model is trained based on the data collected during the first time period in the third data.
[0101] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0102] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0103] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0104] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0105] In a fifth aspect, the present application provides a data processing device, comprising:
[0106] a transceiver module, configured to send first data;
[0107] Receive a first model; the first model is trained based on the first data, and the first model is a base model for time series prediction;
[0108] Sending second data; the second data and the first data are different types of data;
[0109] Receive a second model; the second model is trained based on the second data, and the second model is used to predict data related to the prediction deviation of the base model.
[0110] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0111] In a possible implementation, the first data and the second data include time series data within the same time period.
[0112] In a possible implementation, the transceiver module is further configured to:
[0113] Sending third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0114] Receive a third model; the third model is trained based on the third data, and the third model is used to predict data related to the prediction deviation of the second model.
[0115] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0116] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0117] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0118] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0119] In a possible implementation, the transceiver module is further configured to:
[0120] Sending indication information; the indication information is used to indicate that the first data is used as a training sample of the base model of the time series prediction.
[0121] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0122] In a sixth aspect, the present application provides a data processing device, the device comprising:
[0123] A processing module, configured to obtain first data and second data; the first data and the second data are different types of data collected for the same scene within a first time period;
[0124] Target data is obtained based on the first data and the second data; wherein the target data is a fusion result of the first prediction data and the second prediction data, the first prediction data is obtained by predicting the data at the target time after the first time period based on the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data.
[0125] In a possible implementation, the second prediction data is a prediction residual of the first prediction data determined based on the second data.
[0126] In a possible implementation, the first data and the second data are time series data.
[0127] In a possible implementation, the processing module is specifically configured to:
[0128] Obtaining, based on the first data, predicted data for a target time after the first time period;
[0129] obtaining second prediction data according to the second data;
[0130] Target data is determined according to a fusion result of the first prediction data and the second prediction data.
[0131] In a possible implementation, the processing module is specifically configured to:
[0132] The first data and the second data are sent to a server, and target data sent from the server is received.
[0133] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0134] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0135] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0136] In a possible implementation, the processing module is further configured to:
[0137] Acquire third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0138] The processing module is specifically used to:
[0139] Target data is obtained based on the first data, the second data and the third data; the target data is a fusion result of the first prediction data, the second prediction data and the third prediction data, and the third prediction data is data related to the prediction deviation of the second prediction data determined based on the third data, or the third prediction data is data related to the prediction deviation of the first prediction data determined based on the third data.
[0140] In a possible implementation, the third prediction data is a prediction residual of the first prediction data or the second prediction data determined according to the third data.
[0141] In one possible implementation, the first prediction data is obtained by predicting data at a target time after the first time period based on the first data through a first model, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data through a second model.
[0142] In a possible implementation, the processing module is further used to: before obtaining the first data and the second data, obtain or send fourth data and fifth data to the server; the fourth data and the first data are data of the same category, and the fifth data and the second data are data of the same category; the first model is trained based on the fourth data, and the second model is trained based on the fifth data.
[0143] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0144] In the seventh aspect, an embodiment of the present application provides a data processing device, which may include a memory, a processor, and a bus system, wherein the memory is used to store programs, and the processor is used to execute the programs in the memory to perform the above-mentioned first aspect and any optional method thereof, the above-mentioned second aspect and any optional method thereof, or the above-mentioned third aspect and any optional method thereof.
[0145] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned first aspect and any optional method thereof, the above-mentioned second aspect and any optional method thereof, or the above-mentioned third aspect and any optional method thereof.
[0146] In the ninth aspect, an embodiment of the present application provides a computer program which, when running on a computer, enables the computer to execute the above-mentioned first aspect and any optional method thereof, the above-mentioned second aspect and any optional method thereof, or the above-mentioned third aspect and any optional method thereof.
[0147] In a tenth aspect, the present application provides a chip system comprising a processor configured to support a data processing device in implementing the functions described in the above aspects, such as transmitting or processing data or information described in the above methods. In one possible design, the chip system further comprises a memory configured to store program instructions and data necessary for executing or training the device. The chip system may consist solely of a chip or may include a chip and other discrete components. BRIEF DESCRIPTION OF THE DRAWINGS
[0148] FIG1A is a schematic diagram of a structure of an artificial intelligence main framework;
[0149] 1B and 1C are schematic diagrams of the application system framework of the present invention;
[0150] FIG1D is a schematic diagram of an optional hardware structure of a terminal;
[0151] FIG2 is a schematic diagram of the structure of a server;
[0152] Figures 3 to 5 are schematic diagrams of a system architecture of the present application;
[0153] Figure 6 shows a cloud service process;
[0154] FIG7 is a flowchart of a data processing method provided in an embodiment of the present application;
[0155] 8 to 12 b are flowcharts of a data processing method provided by an embodiment of the present application;
[0156] Figure 12c is a schematic diagram of an interface provided in an embodiment of the present application;
[0157] FIG13 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0158] FIG14 is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application;
[0159] FIG15 is a schematic diagram of the structure of a server provided in an embodiment of the present application;
[0160] FIG16 is a schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0161] The following describes the embodiments of the present invention in conjunction with the accompanying drawings. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.
[0162] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0163] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0164] As used herein, the terms "substantially," "about," and similar terms are used as terms of approximation, not as terms of degree, and are intended to take into account the inherent variations in measurements or calculations that one of ordinary skill in the art would recognize. Furthermore, the use of "may" when describing embodiments of the present invention refers to "one or more possible embodiments." As used herein, the terms "use," "using," and "used" may be considered synonymous with the terms "utilize," "utilizing," and "utilized," respectively. Additionally, the term "exemplary" is intended to refer to an example or illustration.
[0165] First, let's describe the overall workflow of an AI system. See Figure 1A, which shows a schematic diagram of the main AI framework. This AI framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.
[0166] (1) Infrastructure
[0167] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.
[0168] (2) Data
[0169] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0170] (3) Data processing
[0171] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0172] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0173] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0174] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0175] (4) General ability
[0176] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0177] (5) Smart products and industry applications
[0178] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.
[0179] First, we will introduce the application scenarios of this application. This application can be, but is not limited to, applications with time series data prediction functions (hereinafter referred to as prediction applications) or cloud services provided by cloud-side servers. The following are introduced respectively:
[0180] 1. Prediction Applications
[0181] The product form of the embodiment of the present application can be a prediction application. The prediction application can be run on a terminal device or a cloud-side server.
[0182] In one possible implementation, a prediction application can implement a task of predicting time series data based on input data (e.g., time series data), wherein the prediction application can perform the task of predicting time series data in response to the input data (e.g., time series data) to obtain predicted data.
[0183] For example, the above time series data prediction tasks can be, but are not limited to:
[0184] Time series forecasting task: predict the future value of a time series;
[0185] Spatiotemporal forecasting tasks: similar to time series forecasting, but for multiple locations or trajectories;
[0186] Exceedance prediction task: predict whether the upcoming value will exceed a predefined threshold;
[0187] Anomaly detection tasks: timely detection of rare but disruptive events that require action;
[0188] Time series classification task: classify time series into predefined classes;
[0189] Survival analysis task: predicting the time until an event of interest occurs.
[0190] It should be noted that the embodiments of the present application are more inclined to be applied to time series prediction tasks in related fields with layered signal data. The so-called layered signal can be understood as data prediction through multiple different time series related to the data to be predicted.
[0191] For example, in coal mining scenarios (e.g., underground mining of metal or non-metallic ores (typically coal)), gas concentration can be predicted based on time series of gas concentration, wind speed, wind pressure, carbon monoxide, temperature, coal cutting speed, gas extraction volume, or gas extraction pressure. Other gases, such as carbon monoxide and carbon dioxide, can also be predicted.
[0192] For example, in the scenario of factory dust monitoring and prediction, the main cause of dust generation is the use of external force and machinery to process solid materials in industrial production, such as drilling, crushing, cutting, and transportation. Dust density is the main cause of dust explosions. The detection system mainly uses dust online detector sensors to collect dust density in the air. Dust density can be used as the basic layer data, and secondary factors such as air humidity, temperature, and air flow speed are used as sub-priority sets, while machine operation and human operation data serve as another layer of data sets.
[0193] For example, in chemical industry scenarios, such as coal chemical industry, the release of harmful substances can be predicted based on information such as the amount of raw materials (main materials, auxiliary materials) added to the chemical reaction, reaction temperature, time, stirring speed, etc., to predict the concentration of harmful substances.
[0194] For example, in the metallurgical industry scenario, the temperature of the outgoing product can be predicted based on state sensor data (heating power, raw material quantity, raw material temperature, auxiliary material quantity, auxiliary material temperature, reactor wall temperature, etc.).
[0195] For example, in the mineral processing industry, the product grade can be predicted based on the original ore by adjusting equipment parameters to control the crushing particle size, dosage, sedimentation rate and other conditions.
[0196] In addition, the amount of leaked gas can also be predicted in the scenario of underground gas pipeline leakage prediction.
[0197] In one possible implementation, a user can open a prediction application installed on a terminal device and input input data (such as time series data). The prediction application can perform data prediction on the input data using the method provided in the embodiment of the present application and present the predicted data to the user (the presentation method may be, but is not limited to, display, saving, uploading to the cloud, etc.).
[0198] In one possible implementation, a user can open a prediction application installed on a terminal device and enter input data. The prediction application can send the input data to a server on the cloud side. The server on the cloud side performs data prediction on the input data using the method provided in an embodiment of the present application and transmits the predicted data back to the terminal device. The terminal device can present the predicted data to the user (the presentation method can be but is not limited to display, saving, uploading to the cloud side, etc.).
[0199] Next, the prediction application in the embodiment of this application is introduced from the functional architecture and the product architecture that realizes the function.
[0200] Referring to FIG. 1B , FIG. 1B is a schematic diagram of the functional architecture of a prediction application in an embodiment of the present application:
[0201] In one possible implementation, as shown in FIG1B , a prediction application 102 may receive input parameters 101 (e.g., including input data) and generate a prediction result 103. The prediction application 102 may be executed on (for example) at least one computer system and include computer code that, when executed by one or more computers, causes the computers to execute a natural language model trained using the method provided in the embodiments of the present application.
[0202] Referring to FIG. 1C , FIG. 1C is a schematic diagram of the physical architecture for running a prediction application in an embodiment of the present application:
[0203] Referring to FIG1C , FIG1C shows a schematic diagram of a system architecture. The system may include a terminal 100 and a server 200. The server 200 may include one or more servers (FIG1C illustrates one server as an example), and the server 200 may provide a synthesis function service for one or more terminals.
[0204] Among them, the terminal 100 can be installed with a prediction application, or a web page related to the synthesis function can be opened. The above application and web page can provide an interface. The terminal 100 can receive the relevant parameters entered by the user on the synthesis function interface and send the above parameters to the server 200. The server 200 can obtain the processing results based on the received parameters and return the processing results to the terminal 100.
[0205] It should be understood that in some optional implementations, the terminal 100 can also complete the action of obtaining the processing result based on the received parameters by itself without the need for the cooperation of the server, and the embodiments of the present application are not limited to this.
[0206] Next, the product form of the terminal 100 in FIG1C is described;
[0207] The terminal 100 in the embodiment of the present application can be a mobile phone, a tablet computer, a wearable device, an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiment of the present application does not impose any restrictions on this.
[0208] FIG1D shows a schematic diagram of an optional hardware structure of the terminal 100 .
[0209] 1D , the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a processor 170, an external interface 180, and a power supply 190. Those skilled in the art will appreciate that FIG1D is merely an example of a terminal or multi-function device and does not limit the terminal or multi-function device. The terminal or multi-function device may include more or fewer components than shown, or may combine certain components or have different components.
[0210] The input unit 130 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the portable multifunction device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can detect user touch operations on or near it (for example, operations performed on or near the touch screen using a finger, joint, stylus, or any other suitable object) and drive corresponding connected devices according to pre-set programs. The touch screen can detect user touch actions on the touch screen, convert the touch actions into touch signals and transmit them to the processor 170. It can also receive and execute commands sent by the processor 170; the touch signals include at least touch point coordinate information. The touch screen 131 provides an input interface and an output interface between the terminal 100 and the user. Touch screens can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as a volume control button 132 , a switch button 133 , etc.), a trackball, a mouse, a joystick, and the like.
[0211] Among them, the input device 132 can receive input data and the like.
[0212] The display unit 140 may be used to display information input by or provided to the user, various menus of the terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In an embodiment of the present application, the display unit 140 may be used to display the interface of a prediction application, generated prediction data, etc.
[0213] Memory 120 can be used to store instructions and data. It primarily includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files and text. The instruction storage area can store software units such as the operating system, applications, and instructions required for at least one function, or subsets or extensions thereof. It may also include non-volatile random access memory (RAM). It provides processor 170 with management functions for the hardware, software, and data resources within the computing and processing device, supporting control software and applications. It is also used to store multimedia files and running programs and applications.
[0214] The processor 170 is the control center of the terminal 100. It connects all components of the terminal 100 using various interfaces and circuits. By executing instructions stored in the memory 120 and accessing data stored therein, it executes various functions of the terminal 100 and processes data, thereby providing overall control of the terminal device. Optionally, the processor 170 may include one or more processing units. Preferably, the processor 170 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory may be implemented on a single chip; in other embodiments, they may be implemented on separate chips. The processor 170 may also generate corresponding operational control signals and send them to the corresponding components of the computing and processing device. It may also read and process data in the software, particularly the data and programs in the memory 120, to enable the various functional modules therein to perform their corresponding functions, thereby controlling the corresponding components to operate as instructed.
[0215] Among them, the memory 120 can be used to store software codes related to the data processing method, the processor 170 can execute the steps of the chip's data processing method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve corresponding functions.
[0216] The RF unit 110 (optional) can be used to send and receive information or receive and send signals during a call. For example, after receiving downlink information from the base station, it is passed to the processor 170 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF unit 110 can also communicate with network devices and other devices via wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0217] In this embodiment of the present application, the RF unit 110 may send input data to the server 200 and receive prediction data sent by the server 200 .
[0218] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network port.
[0219] The terminal 100 also includes a power supply 190 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.
[0220] The terminal 100 further includes an external interface 180 , which may be a standard Micro USB interface or a multi-pin connector, and may be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100 .
[0221] Although not shown, the terminal 100 may also include a flash, a wireless fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which are not described in detail here. Some or all of the methods described below may be applied to the terminal 100 shown in FIG1D .
[0222] Next, the product form of the server 200 in FIG1C is described;
[0223] FIG2 provides a schematic diagram of the structure of a server 200. As shown in FIG2, the server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other via the bus 201.
[0224] Bus 201 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, FIG2 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.
[0225] The processor 202 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0226] The memory 204 may include volatile memory, such as random access memory (RAM). The memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard drive (HDD), or solid state drive (SSD).
[0227] The memory 204 may be used to store software codes related to the data processing method, and the processor 202 may execute the steps of the data processing method of the chip, and may also schedule other units to implement corresponding functions.
[0228] It should be understood that the above-mentioned terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above-mentioned terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0229] It should be understood that the steps related to the model reasoning process in the embodiments of this application involve AI-related operations. When performing AI operations, the instruction execution architecture of the terminal device and server is not limited to the processor-memory architecture described above. The system architecture provided in the embodiments of this application is described in detail below with reference to Figure 5.
[0230] FIG5 is a schematic diagram of the system architecture provided by an embodiment of the present application. As shown in FIG5 , the system architecture 500 includes an execution device 510 , a training device 520 , a database 530 , a client device 540 , a data storage system 550 , and a data acquisition system 560 .
[0231] The execution device 510 includes a calculation module 511, an I / O interface 512, a pre-processing module 513, and a post-processing module 514. The calculation module 511 may include the target model / rule 501, and the pre-processing module 513 and the post-processing module 514 are optional.
[0232] The execution device 510 may be a terminal device or a server that runs the above-mentioned prediction application.
[0233] The data acquisition device 560 is used to collect training samples. The training samples can be program files (including time series data of multiple categories), etc. After collecting the training samples, the data acquisition device 560 stores them in the database 530.
[0234] The training device 520 can train the neural network based on the training samples maintained in the database 530 to obtain the target model / rule 501.
[0235] It should be noted that, in actual applications, the training samples maintained in the database 530 may not all be collected by the data acquisition device 560, but may also be received from other devices. It should also be noted that the training device 520 may not train the target model / rule 501 entirely based on the training samples maintained in the database 530, but may also obtain training samples from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.
[0236] The target model / rule 501 obtained through training with the training device 520 can be applied to different systems or devices, such as the execution device 510 shown in FIG5 . The execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, an in-vehicle terminal, etc., or a server, etc.
[0237] Specifically, the training device 520 may transfer the trained model to the execution device 510 .
[0238] In Figure 5, the execution device 510 is configured with an input / output (I / O) interface 512 for data interaction with external devices. The user can input data to the I / O interface 512 through the client device 540 (for example, input data in the embodiment of the present application, etc.).
[0239] Preprocessing module 513 and preprocessing module 514 are used to preprocess the input data received by I / O interface 512. It should be understood that preprocessing module 513 and preprocessing module 514 may be absent or only one preprocessing module may be present. If preprocessing module 513 and preprocessing module 514 are absent, computing module 511 may be used directly to process the input data.
[0240] When the execution device 510 preprocesses the input data, or when the computing module 511 of the execution device 510 performs calculations and other related processing, the execution device 510 can call the data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.
[0241] Finally, the I / O interface 512 provides the processing results (such as prediction data, etc.) to the client device 540, and thus provides them to the user.
[0242] In the scenario shown in FIG5 , the user can manually input data, and this "manual input data" can be operated through the interface provided by I / O interface 512. In another scenario, client device 540 can automatically send input data to I / O interface 512. If user authorization is required for client device 540 to automatically send input data, the user can set the corresponding permissions in client device 540. The user can view the results output by execution device 510 on client device 540, and the specific presentation form can be a display, sound, action, or other specific method. Client device 540 can also serve as a data acquisition terminal, collecting input data input into I / O interface 512 and output results from I / O interface 512 as new sample data and storing them in database 530. Of course, collection can also be performed without client device 540, and instead the I / O interface 512 directly stores the input data input into I / O interface 512 and output results from I / O interface 512 as new sample data in database 530.
[0243] It is worth noting that FIG5 is merely a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in FIG5 , the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the execution device 510 can be deployed in the client device 540.
[0244] From the inference side of the model:
[0245] In the embodiment of the present application, the computing module 511 of the above-mentioned execution device 510 can obtain the code stored in the data storage system 550 to implement the steps related to the model reasoning process in the embodiment of the present application.
[0246] In an embodiment of the present application, the computing module 511 of the execution device 510 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0247] Specifically, the computing module 511 of the execution device 510 can be a hardware system with an execution instruction function, and the steps related to the model reasoning process provided in the embodiment of the present application can be software codes stored in the memory. The computing module 511 of the execution device 510 can obtain the software code from the memory and execute the obtained software code to implement the steps related to the model reasoning process provided in the embodiment of the present application.
[0248] It should be understood that the computing module 511 of the execution device 510 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to the model reasoning process provided in the embodiment of the present application can also be implemented by the hardware system that does not have the function of executing instructions in the computing module 511 of the execution device 510, which is not limited here.
[0249] From the training side of the model:
[0250] In an embodiment of the present application, the above-mentioned training device 520 can obtain the code stored in the memory (not shown in Figure 5, which can be integrated into the training device 520 or deployed separately from the training device 520) to implement the steps related to model training in the embodiment of the present application.
[0251] In an embodiment of the present application, the training device 520 may include a hardware circuit (such as an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with an instruction execution function, such as a CPU, DSP, etc., or a hardware system without an instruction execution function, such as an ASIC, FPGA, etc., or a combination of the above-mentioned hardware systems without an instruction execution function and hardware systems with an instruction execution function.
[0252] It should be understood that the training device 520 can be a combination of a hardware system that does not have the function of executing instructions and a hardware system that has the function of executing instructions. Some of the steps related to model training provided in the embodiments of the present application can also be implemented by the hardware system in the training device 520 that does not have the function of executing instructions, which is not limited here.
[0253] 2. Data prediction cloud services provided by the server:
[0254] In a possible implementation, the server may provide data prediction service to the terminal side through an application programming interface (API).
[0255] Among them, the terminal device can send relevant parameters (such as input data) to the server through the API provided by the cloud. The server can obtain processing results (such as predicted data, etc.) based on the received parameters and return the processing results to the terminal.
[0256] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.
[0257] FIG6 shows the process of using a data prediction function cloud service provided by a cloud platform.
[0258] 1. Activate and purchase the forecast service.
[0259] 2. Users can download the software development kit (SDK) corresponding to the prediction service. Cloud platforms usually provide multiple development versions of the SDK for users to choose according to their development environment requirements, such as Java version SDK, Python version SDK, PHP version SDK, Android version SDK, etc.
[0260] 3. After the user downloads the corresponding version of the SDK to the local computer as needed, they import the SDK project into the local development environment, configure and debug it in the local development environment. The local development environment can also be used to develop other functions, forming an application that integrates data prediction functional capabilities.
[0261] 4. When a data prediction application needs to perform synthesis during use, it can trigger an API call for the synthesis function. When the application triggers the synthesis function, it initiates an API request to the running instance of the data prediction service in the cloud environment. The API request carries input data, which is then processed by the running instance in the cloud environment to obtain the processing results (such as predicted data).
[0262] 5. The cloud environment returns the processing results to the application, thereby completing a data prediction function service call.
[0263] 3. Cloud services providing data prediction models provided by the server:
[0264] In one possible implementation, the server can provide data prediction services to the end-side through an application programming interface (API). That is, based on the needs of the end-side, the server can provide the end-side with a time series prediction model service that can meet the needs of the end-side.
[0265] Among them, the terminal device can send relevant parameters (such as input data) to the server through the API provided by the cloud. The server can obtain processing results (such as data prediction models, etc.) based on the received parameters and return the processing results to the terminal.
[0266] The description of the terminal and the server can be the same as that of the above embodiments, and will not be repeated here.
[0267] In addition to applications and cloud services, the implementation form of this application can also be in large-model inference acceleration libraries and large-model application SDKs.
[0268] In order to better understand the solution of the embodiment of the present application, the following takes text generation as an example and briefly introduces the possible application scenarios of the embodiment of the present application in combination with Figures 2 to 4.
[0269] Figure 3 shows a time series data prediction system, which includes user devices and data processing equipment. User devices include smart terminals such as mobile phones, personal computers, or information processing centers. User devices are the initiators of natural language data processing, initiating requests such as language questions and answers or queries. Typically, users initiate requests through their user devices.
[0270] The aforementioned data processing devices can be devices or servers with data processing capabilities, such as cloud servers, network servers, application servers, and management servers. The data processing devices receive query statements, voice, text, and other information from smart terminals via interactive interfaces. They then use their memory and processors to perform language data processing, including machine learning, deep learning, search, reasoning, and decision-making, and then feed the results back to the user device. The memory in a data processing device is a general term encompassing both local storage and databases storing historical data. The databases can be located on the data processing device or on other network servers.
[0271] In the time series data prediction system shown in Figure 3, the user device can receive the user's instructions. For example, the user device can receive a piece of time series data input by the user, and then initiate a request to the data processing device so that the data processing device obtains future data predicted based on the input time series data for the user device.
[0272] In an embodiment of the present application, the user device can receive instructions from the user. For example, the user device can receive a period of time series data input by the user, and then initiate a request to the data processing device, so that the data processing device executes a time series data prediction application for the period of time series data obtained by the user device, thereby obtaining future data predicted based on the input time series data.
[0273] Text In Figure 3, the data processing device can process the above text data using the method provided in the embodiment of the present application.
[0274] Figure 4 shows another time series data prediction system. In Figure 4, the user device directly serves as a data processing device. The user device can directly receive input from the user and process it directly by the hardware of the user device itself. The specific process is similar to that of Figure 3. Please refer to the above description and will not be repeated here.
[0275] The processors in Figures 3 and 4 can perform data training / machine learning / deep learning through a neural network model or other models, and use the model finally trained or learned by the data (such as the first model, second model, third model, etc. in the embodiment of the present application) for time series data to obtain corresponding processing results.
[0276] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.
[0277] (1) Neural Network
[0278] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:
[0279] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0280] (2) Gas: A mixed gas, primarily methane, with some carbon dioxide and nitrogen, that gushes from coal seams is a harmful factor in coal mining operations. It not only pollutes the air and causes suffocation, but also, when the gas content in the air is 5% to 16%, it can cause fires or explosions when exposed to fire, resulting in accidents. Coal is a porous solid, and gas is adsorbed in its pores. When disturbed by mining, it gushes out. The amount of gas gushes out is proportional to factors such as the gas content in the coal seam, the exposed surface area of the coal seam, and the stress of the coal seam (analogous to a sponge soaked in perfume).
[0281] (3) Gas exceeding the limit: The gas concentration in a certain area underground in a coal mine exceeds the prescribed limit. The national ministries and commissions, provinces, cities, counties, coal mining groups and other regulatory departments have clear requirements for the gas concentration limits at different locations underground, generally 0.8% to 1.2%. Usually, the limits of lower-level units are stricter (lower limits) than those of higher-level units.
[0282] (4) Underground coal mine: A coal mine that uses underground mining methods, an underground mining system consisting of a mine shaft and mine tunnels.
[0283] (5) Roadway: Various passages (similar to tunnels) drilled between the surface and the ore body, used for ore transportation, ventilation, drainage, and the passage of personnel and equipment.
[0284] (6) Coal mining working face: In a broad sense, it refers to the working place where minerals or rocks are directly mined. It moves with the progress of mining and is usually surrounded by the coal seam to be mined and a U-shaped (but not limited to U-shaped, this is just an example) tunnel. In a narrow sense, it specifically refers to the tunnel at the bottom of the U-shaped tunnel, where the main coal mining machinery and equipment (coal mining machines, hydraulic supports, etc.) are arranged.
[0285] (7) Coal mining machine: One of the most common coal mining machinery equipment in underground coal mines, it is used to break coal from the coal wall (breaking coal). It will move back and forth within the working surface to break the coal, and its function is similar to the sledgehammer used for demolishing walls during renovation.
[0286] (8) Hydraulic support: One of the most common coal mining machinery and equipment in underground coal mines. It is used to resist the ground pressure and support the space of the coal mining working face. Its function is similar to the pillars supporting the ceiling in a building.
[0287] (9) Centralized control system: This is a centralized control system for underground coal mining machinery and equipment such as shearers and hydraulic supports, used for data access from onboard sensors, status monitoring, and remote control. This is not a mandatory system and is only available in intelligent working faces. The types of monitoring sensors vary from mine to mine, such as only shearer position sensors but no hydraulic support column pressure sensors, or both.
[0288] (10) Backpropagation algorithm
[0289] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.
[0290] (11) Loss function
[0291] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.
[0292] Time series data prediction tasks involve predicting future data based on historical time series data. To achieve more accurate predictions, the input data often consists of multiple categories related to the data being predicted. For example, in coal mining, gas concentration prediction uses data such as gas concentration, wind speed, wind pressure, carbon monoxide, temperature, coal cutting speed, gas extraction volume, or gas extraction pressure as input to predict future gas concentrations.
[0293] Existing prediction tasks are often completed through a single algorithm model. In this approach, the dimensionality of the model input features is fixed, that is, the number of categories of input data is fixed during model training and inference. When the input data is missing, the corresponding dimension has no input, which will cause the algorithm to fail or the prediction accuracy to be low.
[0294] In order to solve the above problems, the present invention provides a data processing method. The model training method of the present invention is described in detail below with reference to the accompanying drawings.
[0295] Refer to Figure 7, which is a flow chart of a data processing method provided in an embodiment of the present application. As shown in Figure 7, a data processing method provided in an embodiment of the present application may include steps 701 to 704, and these steps are described in detail below.
[0296] 701. Obtain first data;
[0297] 702. Acquire second data; the second data and the first data are different types of data;
[0298] The execution entity of step 701 may be a server.
[0299] In one possible implementation, the server may receive first data and second data (for example, but not limited to, sent from the end side), the first data and the second data may be time series data, the first data may be used as training data for the first model, and the second data may be used as training data for the second model.
[0300] The first data and the second data are different types of time series data.
[0301] The first data and the second data may not be time series data, but may be converted into time series data through a certain mapping method.
[0302] In order to make more accurate predictions, the input data is often multiple types of data related to the data to be predicted. For example, when predicting gas concentration in the coal mining field, data such as gas concentration, wind speed, wind pressure, carbon monoxide, temperature, coal cutting speed, gas extraction volume or gas extraction pressure are used as input.
[0303] However, input data is often missing. For example, in a mine environment, due to the instability of the mine environment, sensors and network equipment are greatly affected by environmental factors and often malfunction, resulting in long-term (hourly or daily) missing data for some categories (as shown in Figure 8). This makes the training input data unreliable and affects the prediction accuracy.
[0304] The idea behind this application is to stratify the input data, where each layer of data can include one or more categories of time series data. A corresponding model can be trained for each layer of data, and the final prediction result is based on the fusion of the prediction results of multiple models. This ensures that even if some layers of data are missing, predictions can still be made using the remaining data and the corresponding model, ensuring the prediction effect. In addition, because data from a large number of categories is not used as input to the same model, the algorithm model can still be feasible and effective even if data from some categories is missing.
[0305] In one possible implementation, the true value used in the training of the model corresponding to one or more layers of data can be set to the data to be predicted (for the convenience of description, the model can be called the model for predicting the main body of the data to be predicted, that is, the base model), and the true value used in the training of the model corresponding to the data of other layers can be set to the prediction residual of the above model. The prediction residual can be understood as the prediction result of the difference between the prediction result of the model and the true result.
[0306] In one possible implementation, since the value of the prediction residual is very small compared to the predicted data body, when the input data of the corresponding category of the model used to predict the residual is missing, the impact on the final predicted data is very small. Therefore, the input data of the model used to predict the data body to be predicted can be selected as data with higher reliability (that is, data with a lower possibility of missing data, or data with a higher correlation with the data body to be predicted). When performing model training or inference, since the input data of the model used to predict the data body to be predicted is less likely to be missing, the model for the data body to be predicted is used, thereby improving the prediction accuracy.
[0307] In one possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein the possibility of data missing when collecting the sensor data of the first category is lower than the possibility of data missing when collecting the sensor data of the second category.
[0308] In one possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein the first model is used to predict sensor data of the third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0309] That is, the first data and the second data may be data of different layers as described above. It should be understood that the hierarchical division of data may be preset or specified by the user.
[0310] For example, in a possible implementation, the server may receive indication information sent by the terminal side; the indication information is used to indicate that the first data is used as a training sample of the base model of the time series prediction.
[0311] For example, in online dust detection, the main cause of dust generation is the use of external force and machinery to process solid materials in industrial production, such as drilling, crushing, cutting, and transportation. Dust density is the main cause of dust explosions. The detection system mainly uses the dust online detector sensor to collect dust density in the air. Dust density can be used as the basic layer data, and secondary factors such as air humidity, temperature, and air flow speed are used as the second priority set, while machine operation and human operation data are used as another layer of data set.
[0312] For example, taking the prediction of gas concentration in a mine mining scenario as an example, the signal acquisition equipment involves multiple types of sensors, as shown in Figure 9a. The collected data related to gas concentration, such as but not limited to gas concentration, wind speed, coal mining machine power, hydraulic support pressure, hydraulic support stroke, and other gas-related signals are classified and used as data input for training models at different levels.
[0313] The working face airflow flows unidirectionally in the U-shaped tunnel (other shapes are also available). The main sources of gas outbursts are exposed coal walls, mined coal blocks (fallen coal), lost coal blocks that have been mined but not completely recovered (goaf), and outbursts from adjacent coal seams above and below (adjacent layers). The tunnel is equipped with gas sensors (part of the safety monitoring system, monitoring gas concentration, wind speed, etc., reflecting the concentration distribution of gas in the tunnel space and gas overflow in the goaf), coal mining machine sensors (monitoring the position of the coal mining machine, cutting power / current, onboard gas concentration, etc., reflecting the amount of fallen coal and the exposure speed of fresh coal walls), hydraulic support sensors (monitoring the pressure of the hydraulic support columns, etc., reflecting the gas outburst caused by formation pressure in adjacent layers and coal walls), etc., which are used to monitor the status of gas-related influencing factors.
[0314] Environmental perception mainly relies on the sensors of the above-mentioned safety monitoring system to monitor the gas concentration and wind speed in the tunnel. Equipment monitoring mainly relies on various monitoring sensors on the coal mining equipment, such as the position of the coal mining machine, the cutting power / current of the coal mining machine, the gas concentration on the coal mining machine, the pressure of the hydraulic support column, etc. All the raw data of the sensors are uploaded to the server database in the well computer room through the underground industrial ring network for various application queries.
[0315] Gas concentration data (gas concentration at the upper corner of the working face), shearer data (shearer position, shearer cutting current), and hydraulic support data (hydraulic support column pressure 1# to 183#) are read from the mine database. The data sample is as follows:
[0316] Table 1 Gas concentration in the corner of the working face
[0317] Table 2 Coal mining machine position
[0318] Table 3 Coal mining machine cutting current
[0319] Table 4 Hydraulic support column pressure
[0320] In addition, due to the harsh working environment in coal mines, sensor failures and calibrations are often encountered, and the raw sensor data read from the database often has errors and omissions. Operations such as outlier removal and missing value interpolation are required to clean the data to ensure data quality. Specific operations include: Data cleaning 1: Eliminating abnormal data that does not conform to the actual situation, such as data with gas concentration <0; Data cleaning action 2: Interpolating and resampling various types of data with different sampling rates and missing conditions to make the sampling rate and timestamp consistent, such as processing them into data points with 15s intervals. If there are multiple values within the interval, the maximum value is taken. If a certain type of data is missing for more than a certain time range (such as more than 1 hour), all other types of data within the time range are deleted, that is, the data within the time range is not used (training process). The data sample after cleaning is as follows:
[0321] Table 5
[0322] When stratifying data, groups can be formed based on the quality of each data type (whether large segments are frequently missing, and the missing rate) and the degree of relevance to the prediction target (correlation). Gas concentration data originates from the safety monitoring system, which is subject to mandatory policy requirements. Sensors are well maintained, the data quality is high, and the correlation with the prediction target is the best (autocorrelation). This data can be considered the highest priority data set. Coal mining is one of the main factors in gas emission. Coal mining machine-related data has a high correlation with gas concentration, but sensors are prone to failure, resulting in long periods of data loss. This data set can be considered the second priority data set. Coal seam pressure is also a factor influencing gas emission. Hydraulic support column pressure has a certain correlation with it, but the sensor has a high failure rate and data is easily lost. This data set can be considered the second priority data set. Data related to gas influencing factors include but are not limited to: gas (methane) concentration (with multiple sensors, such as air intake channel, return air channel, working face, return air corner, etc.), wind speed, wind pressure, wind direction, temperature, coal mining machine position / speed, coal mining machine onboard gas concentration, coal mining machine variable frequency output current / power, coal mining machine cutting motor current / power, hydraulic support column pressure, hydraulic support jack stroke, scraper conveyor motor current / power, air door / window opening, coal seam gas content, mining footage, coal seam gas extraction volume, etc.
[0323] Specifically, the missing rate of each data type can be evaluated first. For fixed-sampling-rate data (such as coal mining machines and hydraulic supports), the missing rate = number of missing data points / total number of data points. For non-fixed-sampling-rate data (such as gas concentration), assuming the missing rate of data within the past three months needs to be evaluated, the total number of data points for the past year can be calculated, and the average number of data points for the three months can be obtained as the total number of data points. The missing rate can then be calculated. Data stratification is primarily based on sorting by missing rate, with lower missing rates giving higher priority. If missing rates are similar, the correlation between the data and the prediction target is calculated (e.g., using the mutual information method) and sorted based on the correlation. The data sorting and stratification results in this example are: ① gas concentration data (gas concentration at the corner of the working face), ② coal mining machine data (coal mining machine position, coal mining machine cutting current), and ③ hydraulic support data (hydraulic support column pressure).
[0324] Before inputting the data into the model, you can use the cleaned data to construct various statistical features for the algorithm to use, such as extracting the maximum, minimum, and average values of the data every 5 minutes to form time series features.
[0325] For example, the specific operation can be: construct features using cleaned data. The gas concentration data is divided into short-term, medium-term and long-term trend features. The short-term trend feature is a sliding window with a shorter time range (such as 120s) in the negative direction of the prediction start time t0, and the maximum value and the difference value of adjacent intervals are aggregated at fixed time intervals (such as every 15s); the medium-term trend feature is a sliding window with a moderate time range (such as 5min) in the negative direction of t0, and the maximum value, minimum value, range and mean value in the sliding window are taken; the long-term trend is a sliding window with a longer time range (such as 8h) in the negative direction of t0, and the maximum value in the interval is taken at fixed time intervals (such as every 1h). The characteristic data samples are as follows:
[0326] Table 6
[0327] The characteristics of the shearer position and cutting current data are calculated by sliding a window (e.g., 120 seconds) from the prediction start time t0 to a shorter time range, and taking the average value within a fixed time interval (e.g., every 15 seconds). The characteristic samples are as follows:
[0328] Table 7
[0329] The hydraulic support column pressure data feature is the value at the prediction start time t0. The feature example is as follows:
[0330] Table 8
[0331] 703. Send a first model trained based on the first data to the end side; the first model is a base model for time series prediction;
[0332] Among them, the first model can be a base model for time series prediction, and the first data can include time series data within the first time period and data collected at the target moment after the first time period. When training the first model, the time series data within the first time period can be input into the first model, and the first model can obtain the predicted data at the target moment, and use the data collected at the target moment in the first data as the true value to train the first model.
[0333] 704. Send a second model trained based on the second data to the end side; the second model is used to predict data related to the prediction deviation of the base model.
[0334] Among them, the models corresponding to different layers of data in the embodiments of the present application (for example, the first model, the second model, etc.) can be machine learning or deep learning algorithms, such as ridge regression, neural networks, etc., or they can be large models, and the models of each layer can be different.
[0335] In a possible implementation, the first data and the second data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period.
[0336] When training the second model, the second data can be used as the feedforward input for the pre-trained second model, and the ground truth used in training the second model can be constructed based on the first data and the first model. Since the second model needs to be able to predict data related to the prediction deviation of the base model, the ground truth can be data related to the prediction deviation of the base model.
[0337] Among them, the data related to the prediction deviation here can be the prediction residual, that is, the difference between the prediction result and the true value (or a value close to the true value), or a value that can be mapped to the prediction residual. The data related to the prediction deviation can be data of the same dimension as the output of the base model or data of a different dimension, or it can be an abstract feature representation, which is not limited in this application.
[0338] When constructing the true value used in training the second model, the predicted data for the target moment can be obtained through the first model based on the data collected during the first time period in the first data; the residual between the predicted data and the data collected at the target moment in the first data is used as the first true value, and the first true value can be used as the true value used in training the second model. The data collected during the first time period in the second data can be used as the input for feedforward of the second model before training, and the difference between the predicted result of the second model and the first true value is used to train the second model.
[0339] It should be understood that when performing feedforward of the first model, the output result of the first model whose difference with the corresponding true value is greater than a threshold can be selected. It can be considered that the first model has poor processing accuracy for this part of the input data. Therefore, the second data corresponding to this part of the input data can be used as the data for feedforward of the second model, wherein the so-called "second data corresponding to this part of the input data" can be understood as data aligned in time. For example, through threshold judgment, it can be determined that the first model has poor processing accuracy for the first data (the training samples for training the first model may include other data in addition to the first data). Therefore, data in the same time period (or close to the time period) as the first data can be selected from the training data of the second model as the input for training the second model, that is, prediction deviation is predicted for this part of the data with poor prediction accuracy of the first model and corresponding model training is performed. Therefore, the training cost of the second model can be reduced while ensuring the training accuracy of the second model.
[0340] For example, referring to FIG. 9 b , the first data may be large residual samples selected through the basic prediction results.
[0341] Similarly, a model can be trained to predict data related to the second model's prediction deviation. In one possible implementation, third data can be obtained; the third data, the second data, and the first data are different types of time series data collected for the same scene; a third model trained based on the third data is sent to the client; and the third model is used to predict data related to the second model's prediction deviation.
[0342] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0343] In a possible implementation, the first data, the second data, and the third data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period; based on the data collected within the first time period in the second data, the second model can be used to obtain a prediction residual for the first model; the residual between the prediction residual and the first true value is used as the second true value, and the third model is trained based on the data collected within the first time period in the third data.
[0344] Similarly, in addition to the second model, another model can be trained to predict data related to the prediction deviation of the first model (the type of input data for this model is different from the second data). During inference, the results obtained by this model can be fused with the results obtained by the second model (e.g., weighted average).
[0345] Similarly, in addition to the third model, another model can be trained to predict data related to the prediction deviation of the second model (the input data type of this model is different from the third data). During inference, the results obtained by this model can be fused with the results obtained by the third model (e.g., weighted average).
[0346] In one possible implementation, the features constructed from each partitioned data set serve as inputs to each model layer. Each layer is trained independently. If a new data set is added after training, a new layer of model can be trained. Each layer can output the prediction target, or the first layer can output the basic prediction result, and the remaining layers can output the residual between the prediction result and the target.
[0347] For example, a three-layer model can be constructed based on the data stratification results for training, with the corresponding inputs being gas data, coal mining machine data, and hydraulic support data. The first-layer training sample is the gas data feature combined with the maximum gas concentration within the sliding window range (every 5 minutes) after the t0 timestamp. The example is as follows:
[0348] Table 9
[0349] The first-level base model is trained with samples of all gas data features (normalized). The residuals of each sample are then counted. To ensure the prediction effect of high gas concentration, samples with low prediction values (e.g., δ = -0.05, where 0 is not taken for balanced sampling) are selected. The feature part is replaced with coal mining machine data, and the label part is replaced with the corresponding residual. These are used as training samples for compensation model 1. The sample is as follows:
[0350] Table 10
[0351] Repeat the above process, select samples with low predicted values, replace the feature part with hydraulic support data, and update the residual as the training sample of compensation model 2. The sample is as follows:
[0352] Table 11
[0353] Through the above method, for different levels of signal data, the residual samples and compensation value mechanism are used to construct and train the prediction model in layers, and the prediction results of each layer are integrated to obtain the overall prediction result.
[0354] Refer to Figure 10, which is an illustration of hierarchical construction training. When a new data type is introduced, such as gas extraction data, only one layer of model needs to be added, and samples constructed using the gas extraction data features and new residuals are trained without adjusting the previous model.
[0355] In addition, corresponding to the embodiment of FIG. 7 , from the perspective of the terminal side, the embodiment of the present application further provides a data processing method. Referring to FIG. 11 , the method includes:
[0356] 1101. Send first data;
[0357] 1102. Receive a first model; the first model is trained based on the first data, and the first model is a base model for time series prediction;
[0358] 1103. Send second data; the second data and the first data are different types of data;
[0359] 1104. Receive a second model; the second model is trained based on the second data, and the second model is used to predict data related to the prediction deviation of the base model.
[0360] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0361] In a possible implementation, the first data and the second data include time series data within the same time period.
[0362] In one possible implementation, third data may also be sent; the third data, the second data, and the first data are different types of time series data collected for the same scene; a third model is received; the third model is trained based on the third data, and the third model is used to predict data related to the prediction deviation of the second model.
[0363] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0364] In one possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein, the possibility of data missing when collecting the sensor data of the first category is lower than the possibility of data missing when collecting the sensor data of the second category; or, the first model is used to predict sensor data of the third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0365] In a possible implementation, indication information may also be sent; the indication information is used to indicate that the first data is used as a training sample of a base model for the time series prediction.
[0366] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0367] The above describes the embodiment of the present application from the perspective of model training. Next, a data processing method provided by the present application will be described from the perspective of model reasoning. The model training method of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0368] Referring to Figure 12a, Figure 12a is a flow chart of a data processing method provided in an embodiment of the present application. As shown in Figure 12a, a data processing method provided in an embodiment of the present application may include steps 1201 to 1202, and these steps are described in detail below.
[0369] 1201. Acquire first data and second data; the first data and the second data are different types of data collected for the same scene within a first time period;
[0370] In a possible implementation, the first data and the second data are time series data.
[0371] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0372] In one possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein, the possibility of data missing when collecting the sensor data of the first category is lower than the possibility of data missing when collecting the sensor data of the second category; or, the first model is used to predict sensor data of the third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0373] 1202. Obtain target data based on the first data and the second data; wherein the target data is a fusion result of the first prediction data and the second prediction data, the first prediction data is obtained by predicting the data at the target time after the first time period based on the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data.
[0374] In a possible implementation, the second prediction data is a prediction residual of the first prediction data determined based on the second data.
[0375] In one possible implementation, predicted data for a target time after the first time period can be obtained based on the first data; second predicted data can be obtained based on the second data; and target data can be determined based on a fusion result of the first predicted data and the second predicted data.
[0376] When the second prediction data is a prediction residual of the first prediction data determined according to the second data, the first prediction data and the second prediction data may be added together to obtain a fusion result.
[0377] In one possible implementation, the first data and the second data can be sent to a server. The server can obtain predicted data for a target time after the first time period based on the first data; obtain second predicted data based on the second data; determine target data based on the fusion result of the first predicted data and the second predicted data, and send the target data to the end side. Accordingly, the end side can receive the target data sent from the server.
[0378] In a possible implementation, third data can also be obtained; the third data, the second data and the first data are different types of time series data collected for the same scene; target data can be obtained based on the first data, the second data and the third data; the target data is the fusion result of the first prediction data, the second prediction data and the third prediction data.
[0379] In a possible implementation, the third prediction data is data related to a prediction deviation of the second prediction data determined based on the third data. Furthermore, during fusion, the first prediction data, the second prediction data, and the third prediction data may be fused.
[0380] In one possible implementation, the third prediction data is data related to a prediction deviation of the first prediction data determined based on the third data. Furthermore, during fusion, the second prediction data and the third prediction data may be fused to obtain a fused prediction deviation for the first prediction data, and the fused prediction deviation for the first prediction data may be fused with the first prediction data.
[0381] In a possible implementation, the third prediction data is a prediction residual of the first prediction data or the second prediction data determined according to the third data.
[0382] In one possible implementation, the first prediction data is obtained by predicting data at a target time after the first time period based on the first data through a first model, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data through a second model.
[0383] For the introduction of the first model and the second model, reference can be made to the above embodiments, and the similarities will not be repeated here.
[0384] In one possible implementation, before obtaining the first and second data, fourth and fifth data may also be obtained; the fourth data and the first data are of the same type, and the fifth data and the second data are of the same type; the first model is trained based on the fourth data, and the second model is trained based on the fifth data. In other words, the first model may be trained based on the fourth data, and the second model may be trained based on the fifth data.
[0385] In the embodiments of the present application, after the trained model layers are deployed, they are each provided with the required feature data as input, and independently predicted. The results of each layer are fused to form the final prediction result. If data for a layer is missing, that layer is not included in the fusion. If each layer outputs a prediction target, the fusion method can be a weighted average of the outputs of each layer. For example, if the first layer outputs the basic prediction result and the remaining layers output the residual, the fusion method is linear superposition.
[0386] For example, referring to Figure 12b, which illustrates a model inference architecture, if a sensor fails and the corresponding data is missing, such as a long-term failure of a coal mining machine position sensor, compensation model 1 will not output any results, and the final concentration prediction value will be the output of the base model and compensation model 2. Each layer of the inference process makes independent predictions, and if data is missing at a particular layer, the model at that layer will be disabled.
[0387] After obtaining the target data, the user can be warned based on the target data, for example, through the numerical relationship between the target data and the danger threshold. Referring to Figure 12c, Figure 12c is a front-end interface presented to the user, in which the system can predict the gas concentration curve within the next X minutes and determine the future exceeding limit time based on the set exceeding limit value. At the same time, the warning information is presented in the upper window.
[0388] 13 , which is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in FIG13 , a data processing device 1300 provided in an embodiment of the present application includes:
[0389] The processing module 1301 is configured to obtain first data and second data; the second data and the first data are different types of data;
[0390] For a detailed description of the processing module 1301 , reference may be made to the description of steps 701 and 702 in the above embodiment, and similarities will not be repeated here.
[0391] The transceiver module 1302 is used to send a second model trained based on the second data to the end side; the second model is used to predict data related to the prediction deviation of the base model.
[0392] For a detailed introduction to the transceiver module 1302 , reference may be made to the introduction to steps 703 and 704 in the above embodiment, and similarities will not be repeated here.
[0393] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0394] In a possible implementation, the first data and the second data include time series data within the same time period.
[0395] In a possible implementation, the first data and the second data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period;
[0396] Before sending the second model trained according to the second data to the end side, the processing module is further configured to:
[0397] Obtaining predicted data for the target time using the first model based on the data collected during the first time period in the first data;
[0398] The residual between the predicted data and the data collected at the target time in the first data is used as a first true value, and a second model is trained based on the data collected during the first time period in the second data.
[0399] In a possible implementation, the processing module is further configured to:
[0400] Acquire third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0401] The transceiver module is further used to: send a third model trained based on the third data to the end side; the third model is used to predict data related to the prediction deviation of the second model.
[0402] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0403] In a possible implementation, the first data, the second data, and the third data include data collected for the same scene within a first time period, and the first data also includes data collected at a target time after the first time period;
[0404] The processing module is further configured to obtain, by using the second model, a prediction residual of the first model based on data collected during the first time period in the second data, before sending the third model trained based on the third data to the end side;
[0405] The residual between the prediction residual and the first true value is used as the second true value, and a third model is trained based on the data collected during the first time period in the third data.
[0406] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0407] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0408] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0409] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0410] In addition, an embodiment of the present application provides a data processing device, the device comprising:
[0411] a transceiver module, configured to send first data;
[0412] Receive a first model; the first model is trained based on the first data, and the first model is a base model for time series prediction;
[0413] Sending second data; the second data and the first data are different types of data;
[0414] Receive a second model; the second model is trained based on the second data, and the second model is used to predict data related to the prediction deviation of the base model.
[0415] In one possible implementation, the second model is used to perform residual prediction of the base model.
[0416] In a possible implementation, the first data and the second data include time series data within the same time period.
[0417] In a possible implementation, the transceiver module is further configured to:
[0418] Sending third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0419] Receive a third model; the third model is trained based on the third data, and the third model is used to predict data related to the prediction deviation of the second model.
[0420] In one possible implementation, the third model is used to predict the prediction residual of the prediction residual of the second model.
[0421] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0422] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0423] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0424] In a possible implementation, the transceiver module is further configured to:
[0425] Sending indication information; the indication information is used to indicate that the first data is used as a training sample of the base model of the time series prediction.
[0426] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0427] In addition, an embodiment of the present application further provides a data processing device, the device comprising:
[0428] A processing module, configured to obtain first data and second data; the first data and the second data are different types of data collected for the same scene within a first time period;
[0429] Target data is obtained based on the first data and the second data; wherein the target data is a fusion result of the first prediction data and the second prediction data, the first prediction data is obtained by predicting the data at the target time after the first time period based on the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data.
[0430] In a possible implementation, the second prediction data is a prediction residual of the first prediction data determined based on the second data.
[0431] In a possible implementation, the first data and the second data are time series data.
[0432] In a possible implementation, the processing module is specifically configured to:
[0433] Obtaining, based on the first data, predicted data for a target time after the first time period;
[0434] obtaining second prediction data according to the second data;
[0435] Target data is determined according to a fusion result of the first prediction data and the second prediction data.
[0436] In a possible implementation, the processing module is specifically configured to:
[0437] The first data and the second data are sent to a server, and target data sent from the server is received.
[0438] In a possible implementation, the first data is sensor data of the first category, and the second data is sensor data of the second category; wherein,
[0439] The probability of missing data when collecting the first type of sensor data is lower than the probability of missing data when collecting the second type of sensor data; or
[0440] The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
[0441] In a possible implementation, the processing module is further configured to:
[0442] Acquire third data; the third data, the second data, and the first data are different types of time series data collected for the same scene;
[0443] The processing module is specifically used to:
[0444] Target data is obtained based on the first data, the second data and the third data; the target data is a fusion result of the first prediction data, the second prediction data and the third prediction data, and the third prediction data is data related to the prediction deviation of the second prediction data determined based on the third data, or the third prediction data is data related to the prediction deviation of the first prediction data determined based on the third data.
[0445] In a possible implementation, the third prediction data is a prediction residual of the first prediction data or the second prediction data determined according to the third data.
[0446] In one possible implementation, the first prediction data is obtained by predicting data at a target time after the first time period based on the first data through a first model, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data through a second model.
[0447] In a possible implementation, the processing module is further used to: before obtaining the first data and the second data, obtain or send fourth data and fifth data to the server; the fourth data and the first data are data of the same category, and the fifth data and the second data are data of the same category; the first model is trained based on the fourth data, and the second model is trained based on the fifth data.
[0448] In a possible implementation, the first data and the second data are sensor data related to hazardous substances in a resource collection scenario.
[0449] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 14. Figure 14 is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1400 can be specifically manifested as a virtual reality VR device, a mobile phone, a tablet, a laptop computer, a smart wearable device, a monitoring data processing device or a server, etc., which is not limited here. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403 and a memory 1404 (wherein the number of processors 1403 in the execution device 1400 can be one or more, and Figure 14 takes one processor as an example), wherein the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403 and the memory 1404 may be connected via a bus or other means.
[0450] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0451] Processor 1403 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0452] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1403. Processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1403. The above processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1403 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in memory 1404. Processor 1403 reads information from memory 1404 and, in conjunction with its hardware, completes the steps involved in the model inference process in the above method.
[0453] Receiver 1401 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1402 can be used to output digital or character information through the first interface. Transmitter 1402 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1402 can also include a display device such as a display screen.
[0454] The embodiment of the present application also provides a server. Please refer to Figure 15. Figure 15 is a schematic diagram of the structure of a server provided by the embodiment of the present application. Specifically, the server 1500 is implemented by one or more servers. The server 1500 may have relatively large differences due to different configurations or performance. It may include one or more central processing units (CPUs) 1515 (for example, one or more processors) and memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage medium 1530 can be temporary storage or permanent storage. The program stored in the storage medium 1530 may include one or more modules (not shown in the figure), each module may include a series of instruction operations in the server. Furthermore, the central processing unit 1515 can be configured to communicate with the storage medium 1530 to execute a series of instruction operations in the storage medium 1530 on the server 1500.
[0455] The server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input and output interfaces 1558; or one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0456] In the embodiment of the present application, the central processing unit 1515 is used to execute the data processing method in the above embodiment.
[0457] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0458] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0459] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0460] Specifically, see Figure 16 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor (NPU) 1600. NPU 1600 is mounted on a host CPU (host CPU) as a coprocessor, with tasks assigned by the host CPU. The core of the NPU is arithmetic circuit 1603, which is controlled by controller 1604 to extract matrix data from memory and perform multiplication operations.
[0461] In some implementations, arithmetic circuit 1603 includes multiple processing units (PEs). In some implementations, arithmetic circuit 1603 is a two-dimensional systolic array. Arithmetic circuit 1603 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, arithmetic circuit 1603 is a general-purpose matrix processor.
[0462] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1602 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1601 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1608.
[0463] Unified memory 1606 is used to store input and output data. Weight data is directly transferred to weight memory 1602 through the Direct Memory Access Controller (DMAC) 1605. Input data is also transferred to unified memory 1606 through the DMAC.
[0464] BIU stands for Bus Interface Unit 1610 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1609 .
[0465] The bus interface unit 1610 (BIU) is used for the instruction fetch memory 1609 to obtain instructions from the external memory, and is also used for the storage unit access controller 1605 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0466] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1606 or move weight data to the weight memory 1602 or move input data to the input memory 1601.
[0467] The vector calculation unit 1607 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1603, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0468] In some implementations, vector calculation unit 1607 can store the processed output vector to unified memory 1606. For example, vector calculation unit 1607 can apply a linear function or a nonlinear function to the output of operation circuit 1603, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, vector calculation unit 1607 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to operation circuit 1603, for example, for use in subsequent layers in a neural network.
[0469] An instruction fetch buffer 1609 connected to the controller 1604 is used to store instructions used by the controller 1604;
[0470] Unified memory 1606, input memory 1601, weight memory 1602, and instruction fetch memory 1609 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0471] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0472] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0473] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0474] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0475] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A data processing method, characterized in that, The method includes: Obtain first data; Send to the edge side a first model trained based on the first data; the first model is a base model for time series prediction; Obtain second data; the second data and the first data are different types of data; Send to the edge side a second model trained based on the second data; the second model is used to predict data related to the prediction deviation of the base model.
2. The method according to claim 1, wherein The second model is used to perform residual prediction of the base model.
3. The method according to claim 1 or 2, characterized in that, The first data and the second data include time series data within the same time period.
4. The method according to any one of claims 1 to 3, characterized in that, The first data and the second data include data collected for the same scenario during a first time period, and the first data further includes data collected at a target time after the first time period; Before sending to the edge side the second model trained based on the second data, the method further includes: Based on the data collected during the first time period in the first data, obtain prediction data for the target time through the first model; Use the residual between the prediction data and the data collected at the target time in the first data as the first true value, and train a second model based on the data collected during the first time period in the second data.
5. The method according to claim 4, wherein Training the second model based on the data collected during the first time period in the second data includes: When the first true value is greater than a threshold, train a second model based on the data collected during the first time period in the second data.
6. The method according to any one of claims 1 to 5, characterized in that The method further includes: Obtain third data; the third data, the second data, and the first data are different types of time series data collected for the same scenario; Send to the edge side a third model trained based on the third data; the third model is used to predict data related to the prediction deviation of the second model.
7. The method according to claim 6, wherein The third model is used to predict the prediction residual of the prediction residual of the second model.
8. The method according to claim 6 or 7, characterized in that, The first data, the second data, and the third data include data collected for the same scenario during a first time period, and the first data further includes data collected at a target time after the first time period; Before sending to the edge side the third model trained based on the third data, the method further includes: Based on the data collected during the first time period in the second data, obtain the prediction residual of the first model through the second model; Use the residual between the prediction residual and the first true value as the second true value, and train a third model based on the data collected during the first time period in the third data.
9. The method according to any one of claims 1 to 8, characterized in that The first data is sensor data of a first type, and the second data is sensor data of a second type; wherein, The possibility of data loss when collecting the sensor data of the first type is lower than the possibility of data loss when collecting the sensor data of the second type; or, The first model is used to predict sensor data of a third category based on sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
10. The method according to any one of claims 1 to 9, characterized in that The method further includes: Receiving indication information sent by the terminal side; The indication information is used to indicate that the first data is used as a training sample of the base model for the time series prediction.
11. The method according to any one of claims 1 to 10, characterized in that, The first data and the second data are sensor data related to harmful substances in a resource acquisition scenario.
12. A data processing method, characterized in that, The method includes: Sending the first data; Receiving a first model; the first model is trained according to the first data, and the first model is the base model for time series prediction; Sending the second data; the second data and the first data are data of different categories; Receiving a second model; the second model is trained according to the second data, and the second model is used to predict data related to the prediction deviation of the base model.
13. The method according to claim 12, wherein The second model is used to perform residual prediction of the base model.
14. The method according to claim 12 or 13, wherein The first data and the second data include time series data within the same time period.
15. The method according to any one of claims 12 to 14, characterized in that, The method further includes: Sending third data; the third data, the second data, and the first data are time series data of different categories collected for the same scenario; Receiving a third model; the third model is trained according to the third data, and the third model is used to predict data related to the prediction deviation of the second model.
16. The method according to claim 15, wherein The third model is used to predict the prediction residual of the prediction residual of the second model.
17. The method according to any one of claims 12 to 16, characterized in that, The first data is sensor data of a first category, and the second data is sensor data of a second category; wherein, The possibility of data loss when collecting the sensor data of the first category is lower than the possibility of data loss when collecting the sensor data of the second category; or, The first model is used to predict sensor data of a third category based on the sensor data of the first category, and the correlation between the sensor data of the first category and the sensor data of the third category is higher than the correlation between the sensor data of the second category and the sensor data of the third category.
18. The method according to any one of claims 12 to 17, characterized in that The method further includes: Sending indication information; the indication information is used to indicate that the first data is used as a training sample of the base model for the time series prediction.
19. The method according to any one of claims 12 to 18, characterized in that The first data and the second data are sensor data related to harmful substances in a resource acquisition scenario.
20. A data processing method, characterized in that The method includes: Obtaining first data and second data; the first data and the second data are data of different categories collected for the same scenario within a first time period; According to the first data and the second data, obtaining target data; wherein, the target data is a fusion result of a first prediction data and a second prediction data, the first prediction data is obtained by predicting data at a target moment after the first time period according to the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined according to the second data.
21. The method according to claim 20, wherein The second prediction data is the prediction residual of the first prediction data determined according to the second data.
22. The method according to claim 20 or 21, wherein The first data and the second data are time series data.
23. The method according to any one of claims 20 to 22, characterized in that, Obtaining target data according to the first data and the second data includes: Obtaining prediction data at a target time after the first time period according to the first data; Obtaining second prediction data according to the second data; Determining target data according to the fusion result of the first prediction data and the second prediction data.
24. The method according to any one of claims 20 to 23, characterized in that Obtaining target data according to the first data and the second data includes: Sending the first data and the second data to a server and receiving target data sent from the server.
25. The method according to any one of claims 20 to 24, characterized in that The first data is sensor data of a first type, and the second data is sensor data of a second type; wherein, The possibility of data loss when collecting the sensor data of the first type is lower than the possibility of data loss when collecting the sensor data of the second type; or, The first model is used to predict sensor data of a third type based on the sensor data of the first type, and the correlation between the sensor data of the first type and the sensor data of the third type is higher than the correlation between the sensor data of the second type and the sensor data of the third type.
26. The method according to any one of claims 20 to 25, characterized in that The method further includes: Obtaining third data; the third data, the second data, and the first data are different types of time series data collected for the same scenario; Obtaining target data according to the first data and the second data includes: Obtaining target data according to the first data, the second data, and the third data; the target data is the fusion result of the first prediction data, the second prediction data, and the third prediction data, and the third prediction data is data related to the prediction deviation of the second prediction data determined according to the third data, or the third prediction data is data related to the prediction deviation of the first prediction data determined according to the third data.
27. The method according to claim 26, wherein The third prediction data is the prediction residual of the first prediction data or the second prediction data determined according to the third data.
28. The method according to any one of claims 20 to 27, characterized in that, The first prediction data is obtained by predicting data at the target time after the first time period through a first model according to the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined through a second model according to the second data.
29. The method according to claim 28, wherein Before obtaining the first data and the second data, the method further includes: Obtaining or sending fourth data and fifth data to a server; the fourth data and the first data are of the same type, and the fifth data and the second data are of the same type; the first model is trained according to the fourth data, and the second model is trained according to the fifth data.
30. The method according to any one of claims 20 to 29, characterized in that, The first data and the second data are sensor data related to harmful substances in a resource collection scenario.
31. A data processing device, characterized in that, The device includes: A processing module, configured to obtain first data; obtain second data; the second data and the first data are of different types; A transceiver module, configured to send to the terminal side a first model trained based on the first data; the first model is a base model for time series prediction; and send to the terminal side a second model trained based on the second data; the second model is used to predict data related to the prediction deviation of the base model.
32. The device according to claim 31, wherein, The second model is used to perform residual prediction of the base model.
33. The device according to claim 31 or 32, characterized in that, The first data and the second data include time series data within the same time period.
34. The device according to any one of claims 31 to 33, characterized in that, The first data and the second data include data collected for the same scenario within a first time period, and the first data further includes data collected at a target time after the first time period; Before sending to the terminal side the second model trained based on the second data, the processing module is further configured to: Based on the data collected within the first time period in the first data, obtain, through the first model, prediction data for the target time; Use the residual between the prediction data and the data collected at the target time in the first data as the first true value, and train a second model based on the data collected within the first time period in the second data.
35. A data processing device, characterized in that, The apparatus includes: A transceiver module, configured to send first data; receive a first model; the first model is trained based on the first data and is a base model for time series prediction; send second data; the second data and the first data are of different types; and receive a second model; the second model is trained based on the second data and is used to predict data related to the prediction deviation of the base model.
36. A data processing device, characterized in that, The apparatus includes: A processing module, configured to obtain first data and second data; the first data and the second data are different types of data collected for the same scenario within a first time period; and obtain target data based on the first data and the second data; wherein the target data is a fusion result of first prediction data and second prediction data, the first prediction data is prediction data for the data at a target time after the first time period based on the first data, and the second prediction data is data related to the prediction deviation of the first prediction data determined based on the second data.
37. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers are caused to perform the operations of the method according to any one of claims 1 to 30.
38. A computer program product, characterized in that, Including computer-readable instructions, when the computer-readable instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 30.
39. A system includes at least one processor and at least one memory; the processor and the memory are connected through a communication bus and communicate with each other; The at least one memory is used to store code; The at least one processor is used to execute the code to execute the method according to any one of claims 1 to 30.
40. A chip, characterized in that, Comprising at least one processing unit and an interface circuit, the interface circuit being configured to provide program instructions or data for the at least one processing unit, the at least one processing unit being configured to execute the program instructions to implement the method according to any one of claims 1 to 30.
Citation Information
Patent Citations
Data processing method and device
CN120234607A
Transform-based multi-mode self-correlation compensation time sequence prediction method
CN114580709A
Air compression station mother pipe pressure prediction method and system based on time sequence
CN115544863A
Short-term wind power prediction method and device based on deviation compensation and cascade migration
CN116760010A
Time series data prediction method and system based on time series large model optimization
CN118709153A