Data completion method, model training method, storage medium and program product

By training the data completion model, learning data characteristics and completing it, the problem of low data completion accuracy in the existing technology is solved, and higher data completion accuracy is achieved.

CN120045851APending Publication Date: 2025-05-27BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411866069.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When existing data completion technology faces significant data fluctuations or large amounts of missing data, the completion effect is not ideal and has low accuracy.

Method used

By training the data completion model, learn the characteristics of the data itself, and use the trained model to complete the data. The specific method includes inputting the data to be completed into the data completion model, extracting missing data features at the target position, and completing them based on these features.

Benefits of technology

This method can fully consider the characteristics of the original data and improve the accuracy of data completion, especially when the data fluctuates significantly or the amount of missing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045851A_ABST
    Figure CN120045851A_ABST
Patent Text Reader

Abstract

The invention discloses a data completion method, a model training method, a storage medium and a program product, relates to the technical field of data processing, and can improve the accuracy of data completion. The battery performance detection method comprises the following steps: receiving data to be complemented; missing data at the target position of the to-be-complemented data; inputting the to-be-complemented data into the data complementation model to obtain complemented data; wherein the data completion model is used for extracting data features of missing data of the target position of the to-be-completed data, and completing the missing data of the target position based on the data features to obtain completed data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a data completion method, a model training method, a storage medium, and a program product. Background Art

[0002] Before data analysis or machine learning, in order to improve data quality, the original data is often preprocessed, and data completion is an important step in data preprocessing. During data acquisition or transmission, data loss often occurs due to various possible reasons. In these cases, the missing data needs to be completed to ensure data reliability.

[0003] Existing data completion technologies can meet the completion of data with small fluctuations or small amounts of missing data. However, in the face of significant data fluctuations or large amounts of missing data, the completion effect of existing technologies is not ideal, and there is a problem of low accuracy. Summary of the Invention

[0004] Embodiments of the present application provide a data completion method, a model training method, a storage medium, and a program product for completing missing data. The aim is to solve the problem of low accuracy in data completion in the prior art.

[0005] To achieve the above object, the present application adopts the following technical solutions:

[0006] In a first aspect, a data completion method is provided. The method includes: receiving data to be completed; missing data at a target position of the data to be completed; inputting the data to be completed into a data completion model to obtain completed data; wherein the data completion model is used to extract data features of the missing data at the target position of the data to be completed, and complete the missing data at the target position based on the data features to obtain completed data.

[0007] The data completion method provided by the embodiments of the present application can train a data completion model to learn the characteristics of the data itself, and then use the trained data completion model to complete the data. This method can fully consider the characteristics of the original data itself and improve the accuracy of data completion.

[0008] In some embodiments, the data completion model is an autoencoder with a feature extraction model.

[0009] In some embodiments, the feature extraction model is a transformer model.

[0010] In some embodiments, the data completion model includes an encoder and a decoder. The encoder includes an embedding layer, a feature extraction layer, and a data completion layer. The embedding layer is configured to perform positional encoding on the input data to be completed. The feature extraction layer is configured to infer the data features of the missing data at the target position based on the context information of the data to be completed. The data completion layer is configured to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed. The decoder is configured to perform data reconstruction based on the data features of the missing data at the target position and the complete data to be completed to obtain the completed data.

[0011] In some embodiments, when the data to be completed is time-domain data, before inputting the data to be completed into the data completion model, the method further includes: converting the data to be completed from time-domain data to frequency-domain data.

[0012] In some embodiments, converting the data to be completed from time-domain data to frequency-domain data includes: using the Haar transform to convert the data to be completed from time-domain data to frequency-domain data.

[0013] In some embodiments, using the Haar transform to convert the data to be completed from time-domain data to frequency-domain data includes: performing the Haar transform based on the time-domain data of the data to be completed and the weight value of the time-domain data to obtain the frequency-domain data of the data to be completed; wherein, the weight value of the time-domain data is determined based on the local features of the time-domain data.

[0014] In some embodiments, the local features of the time-domain data include at least one of the following: amplitude, frequency, phase, signal strength.

[0015] In some embodiments, the data to be completed is the battery data of a vehicle.

[0016] In a second aspect, the present application provides a model training method, the method includes: obtaining a plurality of sample data; training an initial model based on the sample data to obtain a trained data completion model; wherein, the data completion model is used to extract the data features of the missing data at the target position of the data to be completed, and complete the missing data at the target position based on the data features to obtain the completed data.

[0017] It can be understood that the training method of the data completion model provided by the present application trains the data completion model by learning the data features of the sample data and completing the missing data based on the data features, which can improve the accuracy of the data completion model.

[0018] In some embodiments, the initial model includes an encoder and a decoder. The encoder includes an embedding layer, a data masking layer, a feature extraction layer, and a data completion layer. The embedding layer is configured to perform positional encoding on the input sample data. The data masking layer is configured to mask the data at the target position of the sample data to obtain data to be completed. The feature extraction layer is configured to infer the data features of the missing data at the target position based on the context information of the data to be completed. The data completion layer is configured to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed. The decoder is configured to perform data reconstruction based on the data features of the missing data at the target position and the complete data to be completed to obtain the completed data.

[0019] In some embodiments, the target position is random.

[0020] In some embodiments, the data masking layer is specifically configured to segment the sample data and select some data from each segment for masking.

[0021] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. The memory stores instructions executable by the processor. When the processor is configured to execute the instructions, the electronic device implements the methods of the first aspect and the second aspect described above.

[0022] In a fourth aspect, the present application provides a computer-readable storage medium, which includes computer software instructions. When the computer software instructions run on an electronic device, the electronic device implements the methods of the first aspect and the second aspect described above.

[0023] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program runs on an electronic device, the electronic device implements the methods of the first aspect and the second aspect described above.

[0024] The beneficial effects of the above aspects to the fifth aspect refer to the corresponding descriptions of the first aspect and the second aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 FIG. is a schematic hardware architecture diagram of a data completion method provided by an embodiment of the present application;

[0027] Figure 2 Flow schematic diagram of a data completion method provided by an embodiment of the present application;

[0028] Figure 3 Schematic diagram of a data completion process provided by an embodiment of the present application;

[0029] Figure 4 Flow schematic diagram of another data completion method provided by an embodiment of the present application;

[0030] Figure 5 Schematic diagram of data time-frequency domain conversion provided by an embodiment of the present application;

[0031] Figure 6 Flow schematic diagram of a training method for a data completion model provided by an embodiment of the present application;

[0032] Figure 7 Schematic diagram of a training process for a data completion model provided by an embodiment of the present application;

[0033] Figure 8 Flow schematic diagram of a data completion method provided by an embodiment of the present application;

[0034] Figure 9 Schematic diagram of the structure of a data completion device provided by an embodiment of the present application;

[0035] Figure 10 Schematic diagram of the structure of a training device for a data completion model provided by an embodiment of the present application;

[0036] Figure 11 Schematic diagram of the structure of an electronic device provided by the present application.

[0037] Reference numerals: acquisition device 101, processing device 102, training device 103. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0039] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "front", "rear", "inner", "outer", etc. is based on the orientation or relative positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application. Without special instructions, under the condition of satisfying the relative positional relationship shown in the drawings, the above-mentioned orientation description can be flexibly set during the actual application process.

[0040] The terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0041] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "mounted", "connected", "coupled", "communicated" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection. It may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0042] In the embodiments of the present application, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, article or device including such element.

[0043] In the embodiments of the present application, words such as "exemplary" or "for example" are used to mean being used as an example, illustration or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0044] In the description of this specification, specific features, structures, materials or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0045] In existing data completion methods, one method reads adjacent normal data from the previous frame or the next frame for completion. Another method selects to query the entire data transmission link in reverse and deletes outliers caused by data transmission to meet data consistency. There is also a method for large segments of missing data. Based on the data closest to the missing data segment, linear or simple interpolation operations are performed to complete the original data.

[0046] However, such a processing flow often has the following disadvantages: For single-frame data loss, by assigning the data values of the frames before and after the missing data, large deviations may occur at long sampling intervals. Even at small sampling intervals, when the data to be completed is rapidly changing data such as vehicle driving data, the completed data may not reflect the true change state of the data, ultimately affecting the completion result and even resulting in errors. For large segments or large ranges of missing data, relying solely on linear interpolation or other basic processing methods may lead to errors in the completed data due to ignoring the characteristics of the data itself.

[0047] To address the above technical problems, this application proposes a data completion method. The idea of this method is as follows: Receive the data to be completed; there is missing data at the target position of the data to be completed; input the data to be completed into the data completion model to obtain the completed data. Among them, the data completion model is used to extract the data characteristics of the missing data at the target position of the data to be completed and complete the missing data at the target position based on the data characteristics to obtain the completed data. The data completion method provided by this application can learn the characteristics of the data itself by training the data completion model, and then use the trained data completion model to complete the data. This method can fully consider the characteristics of the original data itself and improve the accuracy of data completion.

[0048] Figure 1 It is a schematic hardware architecture diagram of a data completion method provided by an embodiment of this application. As Figure 1 shown, it includes: a collection device 101, a processing device 102, and a training device 103.

[0049] The collection device 101 is used to obtain data. Exemplarily, the obtained data can be the battery data of a vehicle.

[0050] In some embodiments, the collection device 101 sends the obtained data to the processing device 102 so that the processing device 102 can complete the data. Among them, the data obtained by the collection device 101 is data with data loss.

[0051] In some embodiments, the acquisition device 101 sends the acquired data to the training device 103 so that the training device 103 trains the data completion model based on the data. Among them, the data acquired by the acquisition device 101 has no data missing situation and can be used as the training sample data of the model.

[0052] Exemplarily, the acquisition device 101 can be a sensor or other device with transceiver functions, and the present application does not limit this.

[0053] The processing device 102 is configured to input the data to be completed into the data completion model to obtain the completed data. Among them, the data completion model is used to extract the data features of the missing data at the target position of the data to be completed, and complete the missing data at the target position based on the data features to obtain the completed data.

[0054] In some embodiments, the processing device 102 is further configured to, when the data to be completed is time-domain data, convert the data to be completed from time-domain data to frequency-domain data before inputting the data to be completed into the data completion model.

[0055] In some embodiments, the processing device 102 is further configured to use the Haar transform to convert the data to be completed from time-domain data to frequency-domain data.

[0056] Exemplarily, the processing device 102 can be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or computer, etc. The specific device form of the processing device 102 in the embodiments of the present application is not limited. Figure 1 In the figure, the processing device 102 is shown as a single server.

[0057] The training device 103 is configured to train the data completion model based on the sample data without data missing situation.

[0058] In some embodiments, the completion model includes a data masking layer for masking the data at the target position of the sample data to simulate the data missing situation.

[0059] In some embodiments, the training device 103 is further configured to compare the value of the loss function between the original sample data and the completed sample data and optimize the data completion model using the value of the loss function.

[0060] Exemplarily, the training device 103 can be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or computer, etc. The specific device form of the training device 103 in the embodiments of the present application is not limited. Figure 1 In the figure, the training device 103 is shown as an example of a single server.

[0061] In some embodiments, the training device 103 may exist independently of the processing device 102; alternatively, the processing device 102 and the training device 103 may be integrated in the same device, and the present application does not limit this. Figure 1 Taking the processing device 102 and the training device 103 as different devices as an example is shown below.

[0062] It should be noted that Figure 1 is only an exemplary framework diagram, Figure 1 the number of devices included therein, the names of each device are not limited, and in addition to Figure 1 the devices shown, other devices may also be included, and the embodiments of the present application do not limit this.

[0063] It should be noted that the application scenarios of the embodiments of the present disclosure are not limited. The system architecture and business scenarios described in the embodiments of the present disclosure are for more clearly explaining the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art know that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0064] The following specifically introduces the data completion method provided by the embodiments of the present application.

[0065] Figure 2 is a schematic flowchart of a data completion method provided by an embodiment of the present application. As Figure 2 shown, it includes:

[0066] S101. Receive the data to be completed.

[0067] In some embodiments, the target position of the data to be completed lacks data.

[0068] In some embodiments, the missing data in the data to be completed may be regularly distributed or randomly distributed, and may be a data segment with a small amount of data or a data segment with a large amount of data. The present application does not limit this.

[0069] In some embodiments, the data to be completed may be data with significant fluctuations or data with minor fluctuations. The present application does not limit this. Exemplarily, the data to be completed may be data of a vehicle's battery (such as a ternary battery), such as battery voltage data, battery temperature data, etc. Or, the data to be completed may also be financial data, such as sales data, user behavior data, etc.

[0070] S102. Input the data to be completed into the data completion model to obtain the completed data.

[0071] In some embodiments, the data completion model is used to extract the data features of the missing data at the target position of the data to be completed, and based on the data features, complete the missing data at the target position to obtain the completed data.

[0072] In some embodiments, the data completion model is an autoencoder with a feature extraction model.

[0073] In some embodiments, the feature extraction model is a Transformer model.

[0074] Among them, the Transformer model is a deep learning architecture based on the self-attention mechanism. In addition to the self-attention mechanism, it also has a multi-head attention mechanism and the ability to process long sequences, and can identify different types of dependencies, especially long-distance dependencies, from multiple different attention sub-controls.

[0075] It can be understood that the data completion method provided by this application can learn the characteristics of the data itself by training the data completion model, and then use the trained data completion model to complete the data. This method can fully consider the characteristics of the original data itself and improve the accuracy of data completion.

[0076] In some embodiments, the data completion model includes an encoder and a decoder. Among them, the encoder includes an embedding layer, a feature extraction layer, and a data completion layer:

[0077] 1. The embedding layer is used to perform position encoding on the input data to be completed.

[0078] In some embodiments, the embedding layer is also used to convert the input data to be completed into an embedding vector.

[0079] Exemplarily, sinusoidal and cosine position encoding can be used to perform position encoding on the input data to be completed, that is, a unique position vector is generated for each position in the data sequence according to different frequencies of the sine and cosine functions; or a rotational position encoding method can be adopted, that is, the position information of the data sequence is embedded into the word vector in the form of a rotation matrix.

[0080] It can be understood that encoding the position can provide necessary position information for data completion to enhance its ability to understand sequence data. By retaining absolute and relative position information, position encoding can help the data completion model identify the context information of the missing data, so as to more accurately predict the missing value.

[0081] 2. The feature extraction layer is used to infer the data features of the missing data at the target position according to the context information of the data to be completed.

[0082] In some embodiments, the feature extraction layer extracts features from the context information of the data to be completed based on a Transformer model.

[0083] Exemplarily, the multi-head attention architecture of the Transformer model is used to complete the feature extraction of the data.

[0084] 3. The data completion layer is used to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed.

[0085] Among them, obtaining the complete data to be completed means that only the data length is completed by the preset constant value, and the reconstruction of the data at the missing position is completed by the decoder.

[0086] Exemplarily, the preset constant value can be selected as 0, the mean value of the data to be completed, the median value of the data to be completed, etc.

[0087] 4. The decoder is used to reconstruct the data based on the data features of the missing data at the target position and the complete data to be completed to obtain the completed data.

[0088] In some embodiments, the decoder adopts a TimeMAE model, and reconstructs the missing data in the data to be completed according to the features of the data to be completed extracted by the feature extraction layer, so as to obtain the completed data.

[0089] In some embodiments, the process of completing the data to be completed includes:

[0090] a1. Obtain the data to be completed and input it into the trained data completion model.

[0091] a2. The embedding layer performs position encoding on the data to be completed.

[0092] a3. The feature extraction layer extracts the context information of the non-missing part to infer the data features of the missing part.

[0093] a4. The data completion layer completes the data to be completed through a preset constant so that the length of the completed data is consistent with the preset data length.

[0094] a5. The decoder reconstructs the data to be completed according to the data features of the missing part, obtains the completed data and outputs it.

[0095] Exemplarily, Figure 3 is a schematic diagram of a data completion process provided by an embodiment of the present application.

[0096] In some embodiments, as Figure 4 shown, before the above step S102, the method further includes:

[0097] S201. When the data to be completed is time-domain data, convert the data to be completed from time-domain data to frequency-domain data.

[0098] In some embodiments, the Haar transform is used to convert the data to be completed from time-domain data to frequency-domain data. Among them, the Haar transform is a simple wavelet transform (also known as the Haar wavelet transform), which transforms the input signal through the Haar basis function. Among them, the Haar basis function is an orthonormal function, which has the characteristics of uniform and rapid convergence. The Haar transform is often used in digital signal processing image compression, especially in the analysis and processing of non-stationary signals.

[0099] A possible implementation is to perform the Haar transform based on the time-domain data of the data to be completed and the weight value of the time-domain data to obtain the frequency-domain data of the data to be completed.

[0100] Exemplarily, using the Haar transform to convert the data to be completed from time-domain data to frequency-domain data can refer to the following formula:

[0101]

[0102] Among them, H k is the converted frequency-domain data, x i is the time-domain data before conversion, h ki is the Haar basis function, w i is the weight value of the time-domain data, and k represents the transformation period of the Haar transform.

[0103] In some embodiments, the weight value w i of the time-domain data is determined based on the local features of the time-domain data.

[0104] In some embodiments, the local features of the time-domain data include at least one of the following: amplitude, frequency, phase, signal strength.

[0105] Exemplarily, when determining the weight value w i of the time-domain data according to the signal strength, the weight value w i of the time-domain data can be determined based on a mapping that can reflect the local average or local maximum of the signal strength of the time-domain data; when determining the weight value w i of the time-domain data according to the frequency, the importance of data with different frequencies can be determined based on the analysis of the jitter components of the signal, and then different weight values w i are assigned to different parts of the data.

[0106] Figure 5 This is a schematic diagram of the time-frequency domain conversion of data provided by the embodiments of the present application. As Figure 5 shown, use the Haar transform to transform a certain feature ( Figure 5Perform transformation on the complete data of the diagonal shaded part) to make it data with more uniform dimensions. In this way, the transformed frequency-domain data will have special meanings in each dimension to represent the characteristic information carried in the data. In addition, as Figure 5 shown, it is also possible to use the inverse Haar transform to achieve the transformation of data from frequency-domain data to time-domain data.

[0107] It can be understood that the above-mentioned Haar transform process with variable weights can adaptively adjust the weight values of local data according to the local characteristics of the data to be completed, so as to capture the important characteristic information in the signal. Compared with the transformation method using fixed weights, it can consider data characteristics more comprehensively and improve the accuracy of data transformation. At the same time, since the data characteristics reflected by the frequency-domain data are richer than those of the time-domain data, converting the time-domain data into frequency-domain data in advance can also prepare for subsequent feature extraction, thereby improving the accuracy of data completion.

[0108] Figure 6 This is a schematic flowchart of a method for training a data completion model provided by an embodiment of the present application. As Figure 6 shown, it includes:

[0109] S301. Obtain a plurality of sample data.

[0110] In some embodiments, the sample data is data without data missing situations.

[0111] In some embodiments, the sample data can be obtained through real-time testing or from a database, and the present application does not make any limitations in this regard.

[0112] In some embodiments, the sample data can be data with significant fluctuations or data with minor fluctuations, and the present application does not make any limitations in this regard. Exemplarily, the sample data can be data of a vehicle's battery (such as a ternary battery), such as battery voltage data, battery temperature data, etc. Or, the data to be completed can also be financial data, such as sales data, user behavior data, etc.

[0113] S302. Train an initial model based on the sample data to obtain a trained data completion model.

[0114] Among them, the data completion model is used to extract the data characteristics of the missing data at the target position of the data to be completed, and complete the missing data at the target position based on the data characteristics to obtain the completed data.

[0115] It can be understood that the method for training a data completion model provided by the present application trains the data completion model by learning the data characteristics of the sample data and completing the missing data based on the data characteristics, which can improve the accuracy of the data completion model.

[0116] In some embodiments, the initial model includes an encoder and a decoder. The encoder includes an embedding layer, a data masking layer, a feature extraction layer, and a data completion layer:

[0117] 1. The embedding layer is used to perform positional encoding on the input sample data.

[0118] 2. The data masking layer is used to mask the data at the target position of the sample data to obtain the data to be completed.

[0119] In some embodiments, the target position is random.

[0120] In some embodiments, the data masking layer is used to segment the sample data, and then, taking each segment of data as the smallest unit, randomly select some data segments for masking, so that the total data volume of the finally masked data segments reaches a preset ratio. For example, 15% or 20% of the sample data is masked. The preset ratio is determined according to the actual situation, and the present application does not limit this.

[0121] In some embodiments, data masking can be achieved by adding random masks to the data, data desensitization, etc.

[0122] It can be understood that by randomly masking the sample data to simulate the data missing situation, the model can learn more general data features, and at the same time increase the diversity of the samples, which helps to improve the accuracy of the model.

[0123] 3. The feature extraction layer is used to infer the data features of the missing data at the target position according to the context information of the data to be completed.

[0124] 4. The data completion layer is used to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed.

[0125] 5. The decoder is used to perform data reconstruction based on the data features of the missing data at the target position and the complete data to be completed to obtain the completed data.

[0126] In some embodiments, when the sample data is time-domain data, before performing positional encoding on the sample data, it further includes converting the sample data from time-domain data to frequency-domain data.

[0127] In some embodiments, after obtaining the completed data, by comparing the sample data with the completed data, the loss value between the sample data and the completed data is obtained; and then the parameters of the feature extraction layer and the decoder of the model are optimized using the loss value. After repeated training, a trained data completion model is obtained.

[0128] In some embodiments, the process of training the data completion model according to the sample data includes:

[0129] b1. Obtain training sample data and input it into the data completion model to be trained.

[0130] b2. The embedding layer performs positional encoding on the training sample data.

[0131] b3. The data masking layer randomly masks the sample data.

[0132] Specifically, randomly masking the sample data includes segmenting the data and then randomly masking the data at a fixed ratio.

[0133] b4. The feature extraction layer extracts the context information of the unmasked part to infer the data features of the masked part.

[0134] b5. The data completion layer completes the sample data through a preset constant so that the length of the completed data is the same as that of the original sample data.

[0135] b6. The decoder reconstructs the sample data according to the data features of the masked part to obtain the completed data and outputs it.

[0136] b7. By comparing the input original sample data and the output completed data, determine the value of the loss function between the two.

[0137] b8. Determine whether the data completion model converges according to the value of the loss function.

[0138] b9. If not, update the parameters of the feature extraction layer and the decoder in the data completion model according to the value of the loss function, and return to step b2.

[0139] b10. If so, determine the currently trained data completion model as the trained data completion model.

[0140] Exemplarily, Figure 7 is a schematic diagram of the training process of a data completion model provided by an embodiment of the present application.

[0141] In some embodiments, the above method for training a data completion model further includes: testing the data completion model trained by the training sample set according to the test sample set, and the process is as follows:

[0142] c1. Obtain the completed data completion model.

[0143] c2. Input the sample data in the test sample set into the trained data completion model to obtain the data completion result in the test sample set.

[0144] c3. Determine the loss value of the test sample set according to the data completion result in the test sample set and the original test sample data.

[0145] c4. Determine whether the model converges according to the loss value of the test sample set and a preset loss value threshold.

[0146] Among them, when the loss value of the test sample set is less than the preset loss value threshold, it is determined that the model converges.

[0147] The above-mentioned preset loss value threshold can be determined by means such as experimental tests, simulation, and expert experience. Exemplarily, the preset loss value threshold can be 0.2.

[0148] c5. If not, then it is determined that the data completion model trained with the training sample set is not successfully trained, and it is necessary to train again according to the training sample set.

[0149] c6. If so, then it is determined that the data completion model trained with the training sample set is successfully trained.

[0150] In some embodiments, when using the trained data completion model, by setting built-in parameters, the model can mask the data masking layer during use, so that the data completion model can normally complete the data completion task.

[0151] It can be understood that the data completion method provided in this application processes the original training sample data in a random masking manner, simulates the situation of data loss, helps the model learn the characteristics of the data itself, enables the model to predict and complete the data at the missing positions according to the characteristics of the data itself, and when the data is time-domain data, first converts the data into frequency-domain data and then trains, which helps the data completion model capture richer data characteristics. Using the data completion model trained in this way for data completion can improve the accuracy of data completion compared with the existing method of simply interpolating for data completion.

[0152] The following introduces the data completion method provided in this application through a complete embodiment, as Figure 8 shown, the method includes the following steps:

[0153] S1. Obtain sample data.

[0154] Among them, the sample data is data without data loss.

[0155] S2. When the sample data is time-domain data, convert the sample data from time-domain data to frequency-domain data.

[0156] In some embodiments, if the sample data is frequency-domain data, step S2 is not executed.

[0157] S3. Perform position encoding on the sample data.

[0158] S4. Randomly mask the sample data.

[0159] In some embodiments, the data masking layer is used to segment the sample data, and then, taking each segment of data as the smallest unit, randomly select some data segments for masking, so that the total data volume of the finally masked data segments reaches a preset ratio, such as masking 15% or 20% of the sample data. The preset ratio is determined according to the actual situation, and the present application does not limit this.

[0160] S5. Extract the features of the sample data.

[0161] In some embodiments, the Transformer model is used to extract the features of the sample data.

[0162] S6. Complete the masked positions of the sample data based on a preset value.

[0163] S7. Reconstruct the sample data based on the extracted features of the sample data to obtain the completed sample data.

[0164] In some embodiments, the TimeMAE decoder is used to reconstruct the sample data.

[0165] S8. Compare the original sample data with the completed sample data to determine the loss value.

[0166] S9. Use the loss value to optimize the model to obtain a trained data completion model.

[0167] In some embodiments, the parameters of the Transformer model and the TimeMAE decoder are optimized using the loss value to obtain a trained data completion model.

[0168] S10. Receive the data to be completed.

[0169] Among them, the data to be completed is data with data missing situations.

[0170] S11. When the data to be completed is time-domain data, convert the data to be completed from time-domain data to frequency-domain data.

[0171] In some embodiments, if the data to be completed is frequency-domain data, step S11 is not executed.

[0172] S12. Input the data to be completed into the trained data completion model to obtain the completed data.

[0173] It can be seen that the above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, the embodiments of the present application provide the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0174] The embodiments of the present application can divide the function modules of the communication scheduling device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules. Optionally, the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0175] Figure 9 It is a schematic structural diagram of a data completion device provided by the embodiments of the present application, which can implement the data completion method provided by the above method embodiments. As Figure 9 shown, the data completion device 300 includes: a communication module 301 and a processing module 302.

[0176] The communication module 301 is used to receive the data to be completed; the missing data at the target position of the data to be completed;

[0177] The processing module 302 is used to receive the data to be completed; the missing data at the target position of the data to be completed.

[0178] A possible implementation manner is that the data completion model is an autoencoder with a feature extraction model.

[0179] A possible implementation manner is that the feature extraction model is a transformer model.

[0180] A possible implementation of the data completion model includes an encoder and a decoder. The encoder includes an embedding layer, a feature extraction layer, and a data completion layer. The embedding layer is used to perform positional encoding on the input data to be completed. The feature extraction layer is used to infer the data features of the missing data at the target position based on the context information of the data to be completed. The data completion layer is used to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed. The decoder is used to perform data reconstruction based on the data features of the missing data at the target position and the complete data to be completed to obtain the completed data.

[0181] In a possible implementation, when the data to be completed is time-domain data, before inputting the data to be completed into the data completion model, the processing module 302 is further configured to convert the data to be completed from time-domain data to frequency-domain data.

[0182] In a possible implementation, the processing module 302 is specifically configured to use the Haar transform to convert the data to be completed from time-domain data to frequency-domain data.

[0183] In a possible implementation, the processing module 301 is specifically configured to perform the Haar transform based on the time-domain data of the data to be completed and the weight value of the time-domain data to obtain the frequency-domain data of the data to be completed. The weight value of the time-domain data is determined based on the local features of the time-domain data.

[0184] In a possible implementation, the local features of the time-domain data include at least one of the following: amplitude, frequency, phase, signal strength.

[0185] In a possible implementation, the data to be completed is the battery data of a vehicle.

[0186] Figure 10 This is a schematic structural diagram of a training device for a data completion model provided by an embodiment of the present application, which can implement the model training method provided by the above method embodiment. As Figure 10 shown, the training device 400 for the data completion model includes: a communication module 401 and a processing module 402.

[0187] The communication module 401 is used to obtain a plurality of sample data;

[0188] The processing module 402 is used to train an initial model based on the sample data to obtain a trained data completion model. The data completion model is used to extract the data features of the missing data at the target position of the data to be completed, and complete the missing data at the target position based on the data features to obtain the completed data.

[0189] A possible implementation manner, the initial model includes an encoder and a decoder. Among them, the encoder includes an embedding layer, a data masking layer, a feature extraction layer, and a data completion layer. The embedding layer is used to perform position encoding on the input sample data. The data masking layer is used to mask the data at the target position of the sample data to obtain data to be completed. The feature extraction layer is used to infer the data features of the missing data at the target position according to the context information of the data to be completed. The data completion layer is used to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed. The decoder is used to perform data reconstruction based on the data features of the missing data at the target position and the complete data to be completed to obtain the completed data.

[0190] A possible implementation manner, the target position is random.

[0191] A possible implementation manner, the data masking layer is specifically used to segment the sample data and select some data from each segment for masking.

[0192] In the case of implementing the functions of the above integrated modules in the form of hardware, an embodiment of the present invention provides a possible structural schematic diagram of the electronic device involved in the above embodiment. As Figure 11 shown, the electronic device 900 includes: a processor 902, a communication interface 903, and a bus 904. Optionally, the electronic device 900 may further include a memory 901.

[0193] The processor 902 may be used to implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 902 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 902 may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0194] The communication interface 903 is used to connect to other devices through a communication network. The communication network may be an Ethernet, a wireless access network, a wireless local area network (WLAN), etc.

[0195] The memory 901 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0196] As a possible implementation, the memory 901 can exist independently of the processor 902. The memory 901 can be connected to the processor 902 through the bus 904 and is used to store instructions or program code. When the processor 902 calls and executes the instructions or program code stored in the memory 901, the data completion method provided by the embodiments of the present invention can be implemented.

[0197] In another possible implementation, the memory 901 can also be integrated with the processor 902.

[0198] The bus 904 can be an extended industry standard architecture (EISA) bus, etc. The bus 904 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the service call device is divided into different functional modules to complete all or part of the functions described above.

[0200] The embodiments of the present application also provide a computer-readable storage medium. All or part of the processes in the above method embodiments can be instructed by computer program instructions to complete the relevant hardware. The program can be stored in the above computer-readable storage medium. When the computer program instructions are executed on the computer, the computer is caused to execute the data completion method described in any one of the above embodiments.

[0201] Exemplarily, the above computer-readable storage medium may include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as Compact Discs (CDs), Digital Versatile Discs (DVDs), etc.), smart cards, and flash memory devices (such as Erasable Programmable Read-Only Memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in this disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data).

[0202] An embodiment of this application also provides a computer program product. The computer program product includes a computer program. When the computer program product runs on a computer, the computer is caused to execute any one of the data completion methods provided in the above embodiments.

[0203] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A data completion method, characterized in that: The method comprises: Receive data to be completed; the target position of the data to be completed is missing data; The data to be completed is input into a data completion model to obtain completed data; wherein the data completion model is used to extract data features of the missing data at the target position of the data to be completed, and complete the missing data at the target position based on the data features to obtain the completed data.

2. The method according to claim 1, characterized in that The data completion model is an autoencoder with a feature extraction model.

3. The method according to claim 2, characterized in that The feature extraction model is a transformer model.

4. The method according to claim 1, characterized in that: The data completion model includes an encoder and a decoder, wherein the encoder includes an embedding layer, a feature extraction layer and a data completion layer; The embedding layer is used to perform position encoding on the input data to be completed; The feature extraction layer is used to infer the data features of the missing data at the target location according to the context information of the data to be completed; The data completion layer is used to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed; The decoder is used to reconstruct data based on the data characteristics of the missing data at the target position and the complete data to be completed to obtain the completed data.

5. The method according to any one of claims 1 to 4, characterized in that: In the case where the data to be completed is time domain data, before inputting the data to be completed into the data completion model, the method further includes: The data to be completed is converted from time domain data to frequency domain data.

6. The method according to claim 5, characterized in that The converting the data to be completed from time domain data to frequency domain data comprises: The data to be completed is converted from time domain data to frequency domain data using Haar transform.

7. The method according to claim 6, characterized in that The method of converting the data to be completed from time domain data to frequency domain data by using Haar transform includes: Based on the time domain data of the data to be completed and the weight value of the time domain data, Haar transform is performed to obtain the frequency domain data of the data to be completed; wherein the weight value of the time domain data is determined based on the local characteristics of the time domain data.

8. The method according to claim 7, characterized in that The local features of the time domain data include at least one of the following: amplitude, frequency, phase, and signal strength.

9. The method according to any one of claims 1 to 8, characterized in that: The data to be completed is battery data of the vehicle.

10. A model training method, characterized in that: The method comprises: Get multiple sample data; The initial model is trained based on the sample data to obtain a trained data completion model; wherein the data completion model is used to extract data features of missing data at a target position of the data to be completed, and the missing data at the target position is completed based on the data features to obtain completed data.

11. The method according to claim 10, characterized in that The initial model includes an encoder and a decoder, wherein the encoder includes an embedding layer, a data covering layer, a feature extraction layer and a data completion layer; The embedding layer is used to perform position encoding on the input sample data; The data covering layer is used to cover the data at the target position of the sample data to obtain the data to be completed; The feature extraction layer is used to infer the data features of the missing data at the target location according to the context information of the data to be completed; The data completion layer is used to fill the missing data at the target position of the data to be completed based on a preset constant value to obtain the complete data to be completed; The decoder is used to reconstruct data based on the data characteristics of the missing data at the target position and the complete data to be completed to obtain the completed data.

12. The method according to claim 11, characterized in that The target location is random.

13. The method according to claim 11, characterized in that The data covering layer is specifically used to segment the sample data and select part of the data from each segment for covering.

14. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable a computer device to implement the method according to any one of claims 1 to 9, or the method according to any one of claims 10 to 13.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium comprises computer-executable instructions, and when the computer-executable instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 9 or the method according to any one of claims 10 to 13.

16. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 9 or the method according to any one of claims 10 to 13.