Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By using adjacent data frames and similarity calculations in the car driving data, and correcting them in combination with vehicle status and charging status, the problems of abnormal, repetition and missing in the car driving data are solved, and the data accuracy and processing efficiency are improved.

CN120045860BActive Publication Date: 2025-07-22CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510511062.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

There are quality problems such as abnormality, duplication and missing in car driving data. The existing preprocessing technology is not comprehensive enough, resulting in low data accuracy and reduced available data, hindering the transformation of data to practical results.

Method used

By acquiring multiple data frames of the target vehicle, using adjacent data frames and similarity calculations to determine the target data frame with high similarity as the data frame at the first moment, correcting it in combination with the vehicle status and charging status, and using an interpolation model to process the missing values, optimizing the data processing flow.

Benefits of technology

It improves the accuracy and reliability of data, reduces the complexity of data processing and hardware resource requirements, and improves the overall quality and processing efficiency of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045860B_ABST
    Figure CN120045860B_ABST
Patent Text Reader

Abstract

The present application relates to a data processing method, apparatus, electronic device and storage medium, and relates to the field of data processing. The method includes: obtaining a first data frame, where the first data frame includes a plurality of data frames collected by a target vehicle at a first moment. Then, based on the adjacent data frames of the first data frame and the first data frame, a second data frame is determined. The adjacent data frames of the first data frame include the data frames collected by the target vehicle at adjacent moments of the first moment. Next, a target data frame with the highest similarity to the second data frame is determined from the first data frame, and the target data frame is used as the data frame at the first moment, which can effectively solve quality problems such as anomalies or repetitions in driving data and obtain high-quality driving data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Automobile driving data covers information such as driving behavior characteristics, energy consumption characteristics, and vehicle component working characteristics. This data not only provides key support for breakthroughs in core technologies such as vehicle performance optimization and battery health management, but also has strategic value for the construction of industrial ecosystems such as charging and swapping infrastructure planning and intelligent transportation system construction.

[0003] However, limited by hardware costs such as in-vehicle sensors, communication network devices, and storage devices, as well as multiple factors such as complex vehicle operating conditions interference, there are generally quality problems such as anomalies, repetitions, and missing values in the time-series driving data of automobiles. These problems seriously interfere with the mining and research of the information contained in the driving data, and greatly hinder the transformation of data into practical results.

[0004] Currently, existing data preprocessing technologies do not comprehensively handle the possible quality defects in driving data. In addition to conventional outliers and missing values, the processing of duplicate data is often ignored. Moreover, when dealing with outliers and missing values, simple deletion methods or smoothing methods are often used, which inevitably cause problems such as a decrease in the amount of available data and low data accuracy. Based on this, it is of great significance to construct a driving data preprocessing method that can comprehensively handle the quality problems of driving data and ensure high data accuracy and small loss of available data volume. Summary of the Invention

[0005] One of the purposes of the present invention is to provide a data processing method, apparatus, electronic device, and storage medium, which can effectively solve quality problems such as anomalies and repetitions in driving data and obtain high-quality driving data.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0007] According to the first aspect provided by the present invention, a data processing method is provided. The method includes: obtaining a first data frame, where the first data frame includes multiple data frames collected by a target vehicle at a first moment. Then, based on the adjacent data frames of the first data frame and the first data frame, a second data frame is determined. The adjacent data frames of the first data frame include the data frames collected by the target vehicle at adjacent moments of the first moment. Then, the target data frame with the highest similarity to the second data frame is determined from the first data frame, and the target data frame is used as the data frame at the first moment.

[0008] According to the above technical means, when there are data frames with the same acquisition time in the data of the target vehicle, the second data frame can be determined based on the data frames with the same acquisition time (i.e., the first data frame) and the adjacent data frames of the first data frame. The second data frame obtained in this way can comprehensively reflect the running information of the target vehicle at multiple moments. Then, the target data frame with the highest similarity to the second data frame in the first data frame is used as the data frame at the first moment, which can effectively remove the data frames with low similarity in the first data frame. In this way, not only can the accuracy and reliability of the finally retained target data frame be ensured, but also the overall quality of the vehicle data can be improved.

[0009] In a possible way, to determine the second data frame based on the adjacent data frames of the first data frame and the first data frame, it may specifically include: when the parameters included in multiple data frames meet the first judgment condition, determining the second data frame based on the adjacent data frames of the first data frame and the first data frame. Each data frame includes continuous parameters and / or discrete parameters.

[0010] According to the above technical means, the present application can perform the subsequent operation of determining the second data frame only when the parameters in multiple data frames meet the first judgment condition, avoiding the ineffective processing of data that does not meet specific requirements, and greatly improving the accuracy of data processing.

[0011] In a possible way, the first judgment condition includes: the range of at least one continuous parameter included in multiple data frames is greater than the range threshold; and / or, the discrete parameters included in multiple data frames are different.

[0012] According to the above technical means, the present application can quickly determine the data frames that meet the first judgment condition based on the above first judgment condition.

[0013] In a possible way, the above method further includes: when the parameters included in multiple data frames meet the second judgment condition, selecting any data frame from the first data frame as the data frame at the first moment.

[0014] According to the above technical means, when the parameters of multiple data frames meet the second judgment condition, the present application can directly select any one from the first data frame as the data frame at the first moment, without performing complex operations such as similarity calculation and multi-data frame correlation analysis, reducing the demand for hardware computing resources.

[0015] In a possible way, the second judgment condition includes: the range of each continuous parameter included in multiple data frames is less than or equal to the range threshold, and each discrete parameter included in multiple data frames is the same.

[0016] According to the above technical means, the present application can quickly determine the data frames that meet the second judgment condition based on the above second judgment condition.

[0017] In a possible way, the above method further includes: when the number of adjacent data frames of the first data frame is less than a threshold, generating a third data frame based on the adjacent data frames of the first data frame.

[0018] On this basis, determining the second data frame based on the adjacent data frames of the first data frame and the first data frame may specifically include: determining the second data frame based on the adjacent data frames of the first data frame, the first data frame, and the third data frame.

[0019] According to the above technical means, when the number of adjacent data frames of the first data frame is insufficient, the present application can generate a third data frame in combination with the adjacent data frames of the first data frame. In the subsequent process of determining the second data frame, the third data frame is used as supplementary information and combined with the adjacent data frames of the first data frame and the first data frame, further optimizing the process of determining the second data frame, enabling the second data frame to integrate more dimensional information, and improving the overall efficiency of data utilization.

[0020] In a possible way, determining the target data frame with the highest similarity to the second data frame from the first data frame may specifically include: for each data frame in the first data frame, determining the degree of difference between the data frame and the second data frame, and the degree of difference is negatively correlated with the similarity. Then, the data frame with the smallest degree of difference between the first data frame and the second data frame is used as the target data frame.

[0021] In a possible way, the above method further includes: obtaining a plurality of data frames collected by the target vehicle at different times. Then, for each data frame in the plurality of data frames collected by the target vehicle at different times, determining the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame. Finally, for any two adjacent data frames in the plurality of data frames collected by the target vehicle at different times, determining the data frames belonging to the same driving segment based on the vehicle running states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames. Wherein, the vehicle state refers to the state of the vehicle in different working modes, and the vehicle running state includes different states during the running process of the vehicle.

[0022] On this basis, obtaining the first data frame may specifically include: obtaining the first data frame from the plurality of data frames included in the same driving segment.

[0023] According to the above technical means, the present application can determine the vehicle operating state by comprehensively considering the vehicle state and charging state corresponding to each data frame. This method can more comprehensively and accurately reflect the actual operating conditions of the vehicle at different times. Subsequently, based on the vehicle operating states and time intervals corresponding to any two adjacent data frames, it is determined whether they belong to the same driving segment. This method can reasonably divide the driving process of the vehicle, thereby obtaining data with clear travel boundaries. At the same time, accurate driving segment division is of great significance for analyzing aspects such as the driving behavior, energy consumption, and driving route of the vehicle. In addition, obtaining the first data frame from multiple data frames included in the same driving segment makes the data processing more targeted and can improve the data processing efficiency.

[0024] In a possible way, the vehicle operating state includes a first operating state and a second operating state. The first operating state means that the vehicle state is the starting state, and the charging state is the driving charging state or the non-charging state. The driving charging state means that during the vehicle driving process, the battery is in the energy recovery state. The second operating state means the state when the vehicle state corresponding to the data frame is the non-starting state.

[0025] In a possible way, before determining the vehicle operating state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame, the above method further includes: for each data frame included in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meet the preset correction conditions, based on the charging state, vehicle speed, and total current corresponding to each data frame, correct the vehicle state and / or charging state corresponding to each data frame.

[0026] According to the above technical means, the present application can perform correction based on multi-dimensional information such as the charging state, vehicle speed, and total current when the vehicle state and / or charging state meet the preset correction conditions, and can correct data gaps, anomalies, and inaccuracies caused by various factors (such as sensor errors, signal interference, etc.).

[0027] In a possible way, the correction conditions include any one of the following: the vehicle state corresponding to each data frame included in the same driving segment is other vehicle states, and the other vehicle states refer to vehicle states other than the starting state and the extinguished state. The vehicle state corresponding to each data frame included in the same driving segment is the starting state and the charging state is other charging states, and the other charging states refer to charging states other than the driving charging state and the non-charging state.

[0028] According to the above technical means, the present application can specifically screen out the data frames that need to be corrected based on the above preset correction conditions, which can improve the data processing efficiency and save computing resources and time costs.

[0029] In one possible way, the above method further includes: for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, successively using the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained. The missing value includes a parameter with a null value or a parameter with a value exceeding a preset range.

[0030] Among them, the order of the sub-imputation models for performing imputation processing on the missing value is related to the accuracy index of the sub-imputation model. The accuracy index is used to characterize the accuracy rate of the imputation result obtained by using the sub-imputation model.

[0031] According to the above technical means, the present application can determine the imputation order according to the accuracy index of the sub-imputation model, and preferentially use the sub-imputation model with higher accuracy for imputation, so as to ensure the accuracy rate of the imputation result to the greatest extent. In addition, compared with directly deleting the data frame containing the missing value, using the imputation model for processing can retain more valid data and reduce data waste.

[0032] According to the second aspect provided by the present invention, a data processing device is provided. The device includes: an acquisition unit, a first determination unit, a second determination unit, and a processing unit.

[0033] The acquisition unit is configured to acquire a first data frame, and the first data frame includes a plurality of data frames collected by the target vehicle at a first moment.

[0034] The first determination unit is configured to determine a second data frame based on the adjacent data frames of the first data frame and the first data frame. The adjacent data frames of the first data frame include the data frames collected by the target vehicle at adjacent moments of the first moment.

[0035] The second determination unit is configured to determine a target data frame with the highest similarity to the second data frame from the first data frame.

[0036] The processing unit is configured to use the target data frame as the data frame at the first moment.

[0037] In one possible way, the first determination unit is further configured to determine the second data frame based on the adjacent data frames of the first data frame and the first data frame when the parameters included in the plurality of data frames meet the first judgment condition.

[0038] In one possible way, the first determination unit is further configured to generate a third data frame based on the adjacent data frames of the first data frame when the number of the adjacent data frames of the first data frame is less than a threshold. Then, based on the adjacent data frames of the first data frame, the first data frame, and the third data frame, the second data frame is determined.

[0039] In a possible way, the second determination unit is configured to determine, for each data frame in the first data frame, the degree of difference between the data frame and the second data frame, and use the data frame with the smallest degree of difference between the first data frame and the second data frame as the target data frame.

[0040] In a possible way, the acquisition unit is further configured to acquire a plurality of data frames collected by the target vehicle at different times.

[0041] In a possible way, the first determination unit is further configured to determine, for each data frame in the plurality of data frames collected by the target vehicle at different times, the vehicle operating state corresponding to the data frame based on the vehicle state and the charging state corresponding to the data frame.

[0042] In a possible way, the first determination unit is further configured to determine, for any two adjacent data frames in the plurality of data frames collected by the target vehicle at different times, the data frames belonging to the same driving segment based on the vehicle operating states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames.

[0043] In a possible way, the acquisition unit is further configured to acquire the first data frame from the plurality of data frames included in the same driving segment.

[0044] In a possible way, the processing unit is further configured to, for each data frame included in the same driving segment, when the vehicle state and / or the charging state corresponding to each data frame meet the preset correction conditions, correct the vehicle state and / or the charging state corresponding to each data frame based on the charging state, the vehicle speed, and the total current corresponding to each data frame.

[0045] In a possible way, the processing unit is further configured to, for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, sequentially use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained.

[0046] According to the third aspect of the present invention, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement the method according to the first aspect and any possible implementation manner thereof.

[0047] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the processing device, enabling the processing device to execute the method according to the first aspect and any possible implementation manner thereof.

[0048] According to the fifth aspect provided by the present invention, there is provided a computer program product, which includes computer instructions. When the computer instructions run on a processing device, the processing device is caused to execute the method according to the first aspect and any possible implementation manner thereof as described above.

[0049] Therefore, the above technical features of the present invention have the following beneficial effects:

[0050] (1) When there are data frames with the same acquisition time in the data of the target vehicle, according to the data frames with the same acquisition time (i.e., the first data frames) and the adjacent data frames of the first data frames, the second data frames are determined. The second data frames obtained in this way can comprehensively reflect the running information of the target vehicle at multiple moments. Then, the target data frame with the highest similarity to the second data frame in the first data frame is used as the data frame at the first moment, which can effectively remove the data frames with low similarity in the first data frame. In this way, not only can the accuracy and reliability of the finally retained target data frames be ensured, but also the overall quality of the vehicle data can be improved.

[0051] (2) The subsequent operation of determining the second data frames can be performed only when the parameters in multiple data frames meet the first judgment condition, avoiding the ineffective processing of data that does not meet specific requirements and greatly improving the accuracy of data processing.

[0052] (3) Based on the above first judgment condition, the data frames that meet the first judgment condition can be quickly determined.

[0053] (4) When the parameters of multiple data frames meet the second judgment condition, any one of the first data frames can be directly selected as the data frame at the first moment without performing complex operations such as similarity calculation and multi-data frame correlation analysis, reducing the demand for hardware computing resources.

[0054] (5) Based on the above second judgment condition, the data frames that meet the second judgment condition can be quickly determined.

[0055] (6) When the number of adjacent data frames of the first data frame is insufficient, the third data frames are generated by combining the adjacent data frames of the first data frame. In the subsequent process of determining the second data frames, the third data frames are used as supplementary information and combined with the adjacent data frames and the first data frame of the first data frame, further optimizing the process of determining the second data frames, enabling the second data frames to integrate more dimensional information and improving the overall efficiency of data utilization.

[0056] (7) The vehicle operating state can be determined by comprehensively considering the vehicle state and charging state corresponding to each data frame. This method can more comprehensively and accurately reflect the actual operating conditions of the vehicle at different times. Subsequently, based on the vehicle operating states and time intervals corresponding to any two adjacent data frames, it is determined whether they belong to the same driving segment. This method can reasonably divide the driving process of the vehicle, thereby obtaining data with clear travel boundaries. At the same time, accurate driving segment division is of great significance for analyzing aspects such as the driving behavior, energy consumption, and driving route of the vehicle. In addition, obtaining the first data frame from multiple data frames included in the same driving segment makes the data processing more targeted and can improve the data processing efficiency.

[0057] (8) When the vehicle state and / or charging state meet the preset correction conditions, correction can be performed based on multi-dimensional information such as the charging state, vehicle speed, and total current, which can correct data gaps, anomalies, and inaccuracies caused by various factors (such as sensor errors, signal interference, etc.).

[0058] (9) Based on the above preset correction conditions, data frames that need to be corrected can be selectively screened, which can improve the efficiency of data processing and save computing resources and time costs.

[0059] (10) The interpolation order can be determined according to the accuracy index of the sub-interpolation model, and the sub-interpolation model with higher accuracy is preferentially used for interpolation, which can maximize the accuracy of the interpolation result. In addition, compared with directly deleting data frames containing missing values, using the interpolation model for processing can retain more valid data and reduce data waste. Description of the Drawings

[0060] Figure 1 It is an architecture diagram of a data processing system provided by an embodiment of the present application;

[0061] Figure 2 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 1 ;

[0062] Figure 3 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;

[0063] Figure 4 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 3 ;

[0064] Figure 5 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 4 ;

[0065] Figure 6 Flow schematic of a data processing method provided by an embodiment of the present application Figure 5 ;

[0066] Figure 7 Flow schematic of a data processing method provided by an embodiment of the present application Figure 6 ;

[0067] Figure 8 Flow schematic of a data processing method provided by an embodiment of the present application Figure 7 ;

[0068] Figure 9 Flow schematic of a data processing method provided by an embodiment of the present application Figure 8 ;

[0069] Figure 10 Flow schematic diagram of a method for constructing a target integrated interpolation model provided by an embodiment of the present application;

[0070] Figure 11 Flow schematic of a data processing method provided by an embodiment of the present application Figure 9 ;

[0071] Figure 12 Structure schematic diagram of a data processing device provided by an embodiment of the present application;

[0072] Figure 13 Block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0073] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0074] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0075] In the embodiments of the present application, words such as "exemplary", "for example", or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary", "for example", or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example", or "for example" is intended to present related concepts in a specific way.

[0076] First, the relevant technologies involved in this application are explained to facilitate understanding by those skilled in the art.

[0077] Driven by the global energy structure transformation and the need for low-carbon development, electric vehicles have become the strategic direction of the automotive industry innovation with their core advantages of high energy efficiency and low emissions. With the continuous breakthroughs in battery, motor, and electronic control technologies and the in-depth application of intelligent network technology, the real-time operation data of electric vehicles has shown an explosive growth trend. The vehicle supervision big data platform built in accordance with relevant technical specifications has effectively opened up the collection, transmission and storage links of vehicle driving data, forming a massive data set covering driving behavior characteristics, energy consumption characteristics, and vehicle component working characteristics. These data not only provide key support for core technology breakthroughs such as vehicle performance optimization and battery health management, but also have strategic value for the construction of industrial ecology such as charging and swapping infrastructure planning and intelligent transportation system construction.

[0078] However, due to the cost constraints of on-board sensors, communication network equipment, storage devices and other hardware, as well as the influence of multiple factors such as interference from complex vehicle working conditions, the original electric vehicle time series driving data generally has quality problems such as anomalies, duplications, and missing data, and there is a lack of clear boundaries for each trip data. These problems seriously interfere with the mining of information contained in driving data and greatly hinder the transformation of data into practical results.

[0079] At present, there are many deficiencies in the relevant technologies related to vehicle driving data preprocessing. On the one hand, the existing preprocessing solutions are mostly developed for the private protocol data of specific automobile companies, which are difficult to apply to driving data collected in accordance with the common standard GB / T32960 and lack universality; on the other hand, the processing of possible quality defects in driving data is not comprehensive enough, and in addition to conventional outliers and missing values, the processing of duplicate data is often ignored. In addition, when processing outliers and missing values, simple deletion methods or time series interpolation and smoothing methods are often used, which inevitably leads to problems such as a decrease in the amount of available data and low data accuracy.

[0080] Based on this, there is an urgent need to build a driving data preprocessing method that is highly versatile, can comprehensively deal with driving data quality issues, and ensure high data accuracy and small loss of available data volume, so as to build a high-quality data base for automobile-related research and ensure the authenticity, effectiveness and stability of its results.

[0081] To solve the above technical problems, an embodiment of the present application provides a data processing method. When there are data frames with the same acquisition time in the data of the target vehicle, the second data frame can be determined according to the data frames with the same acquisition time (i.e., the first data frame) and the adjacent data frames of the first data frame. The second data frame obtained in this way can comprehensively reflect the running information of the target vehicle at multiple moments. Then, the target data frame with the highest similarity to the second data frame in the first data frame is used as the data frame at the first moment, which can effectively remove the data frames with low similarity in the first data frame. In this way, not only can the accuracy and reliability of the finally retained target data frame be ensured, but also the overall quality of the vehicle data can be improved.

[0082] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0083] Figure 1 It is an architecture diagram of a data processing system provided by an embodiment of the present application, as Figure 1 shown. The system architecture includes: Server 101.

[0084] Among them, Server 101 can be a high-performance server that provides various services on the Internet. It can be an independent physical server, or a server cluster composed of multiple physical servers, or at least one of the cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data or artificial intelligence platforms. The embodiments of the present application do not limit this. Of course, the server can also include other functions to provide more comprehensive and diverse services.

[0085] The Server 101 in the embodiments of the present application can be a single server, a server cluster, or a cloud server. The embodiments of the present application do not limit this.

[0086] In the embodiments of the present application, when there are multiple data frames collected at the same moment in the data of the target vehicle, Server 101 can obtain the first data frame, where the first data frame includes multiple data frames collected by the target vehicle at the first moment. Then, Server 101 can determine the second data frame based on the adjacent data frames of the first data frame and the first data frame, and determine the similarity between each data frame in the first data frame and the second data frame. Then, Server 101 can use the data frame with the highest similarity to the second data frame in the first data frame as the data frame at the first moment.

[0087] The data processing system provided by the embodiments of the present application can be configured in a vehicle. A vehicle can also be referred to as a transportation vehicle (vehicle), a mobile carrier, an electric vehicle (EV), a hybrid electric vehicle (HEV), a plug-in hybrid electric vehicle (PHEV), a fuel cell vehicle (FCV), an autonomous vehicle, an intelligent and connected vehicle (ICV), a driverless vehicle, etc.

[0088] In the embodiments of the present application, the vehicle can be a sedan, a sport utility vehicle (SUV), a truck, an electric vehicle, a motorcycle, a tricycle, a special vehicle (such as an ambulance, a fire truck, a police car, etc.), a driverless taxi, an intelligent and connected bus, an autonomous logistics vehicle, an electric truck, etc. In addition, this method is also applicable to various special vehicles, such as agricultural vehicles, mining vehicles, forestry vehicles, airport vehicles, port vehicles, etc. The present application does not make specific limitations in this regard.

[0089] For the convenience of understanding, the data processing method provided by the present application will be specifically introduced below in conjunction with the accompanying drawings.

[0090] Figure 2 It is a schematic flowchart of a data processing method provided by the embodiments of the present application. As Figure 2 shown, this method is executed by the Figure 1 server shown, and this method includes:

[0091] S201. Obtain a first data frame.

[0092] Among them, the first data frame includes multiple data frames collected by the target vehicle at the first moment, that is, the collection times of the multiple data frames are the same.

[0093] In some embodiments, the data frames collected by the target vehicle at different moments are stored in the server. On this basis, the server can screen out multiple data frames collected at the first moment from the data frames collected at different moments, so as to obtain the first data frame including multiple collection times that are the same.

[0094] S202. Determine a second data frame based on the adjacent data frames and the first data frame of the first data frame.

[0095] Among them, the adjacent data frames of the first data frame include the data frames collected by the target vehicle at adjacent moments of the first moment. In the embodiments of the present application, the adjacent moments of the first moment may be moments adjacent to or close to the first moment. For example, if the first moment is 10 o'clock, the adjacent moments of the first moment may be 10:01 or 10:05, and there is no limitation thereto.

[0096] In some embodiments, the server may determine the adjacent data frames of the first data frame based on the first moment and the time domain radius threshold. After that, the server may add the first data frame and the adjacent data frames of the first data frame to the domain data frame cluster (i.e., the data frame set), and determine the vector corresponding to each data frame in the domain data frame cluster. Then, the server may determine the second data frame based on the number of data frames included in the domain data frame cluster and the vector corresponding to each data frame.

[0097] In the embodiments of the present application, the vector corresponding to the data frame is the vector composed of the parameters included in the data frame.

[0098] Exemplarily, taking the first moment as t and the time domain radius threshold as as an example. The server adds the data frames with the acquisition moment in the range to the domain data frame cluster, and the other data frames in the domain data frame cluster except the first data frame are the adjacent data frames of the first data frame. After that, the server may determine the second data frame with reference to the following formula 1.

[0099] (Formula 1).

[0100] Among them, is the second data frame (which may also be referred to as the center of the domain data frame cluster), is the number of data frames included in the domain data frame cluster, is the vector corresponding to the i-th data frame in the domain data frame cluster.

[0101] S203. Determine the target data frame with the highest similarity to the second data frame from the first data frame.

[0102] In some embodiments, for each data frame in the first data frame, the server may determine the degree of difference between the data frame and the second data frame. After that, the server may use the data frame with the smallest degree of difference between the first data frame and the second data frame as the target data frame.

[0103] Among them, the degree of difference is negatively correlated with the similarity. For example, the smaller the degree of difference between any data frame in the first data frame and the second data frame, the higher the similarity between the data frame and the second data frame, and the larger the degree of difference between any data frame in the first data frame and the second data frame, the lower the similarity between the data frame and the second data frame.

[0104] The embodiments of the present application do not limit the way of characterizing the degree of difference. For example, the Minkowski distance (such as the Euclidean distance, Manhattan distance, etc.) can be used to characterize the degree of difference, or the Mahalanobis distance can be used to characterize the degree of difference, which is not limited herein. Among them, the distance is positively correlated with the degree of difference, that is, the smaller the distance, the smaller the degree of difference, and the larger the distance, the larger the degree of difference.

[0105] Exemplarily, taking the use of the Mahalanobis distance to characterize the degree of difference as an example. The server can calculate the Mahalanobis distance between each data frame in the first data frame and the second data frame, and then use the data frame in the first data frame with the smallest Mahalanobis distance from the second data frame as the target data frame. Among them, the calculation formula of the Mahalanobis distance can refer to the following formula 2.

[0106] (Formula 2).

[0107] Among them, is the Mahalanobis distance from the data frame to the second data frame, is the vector corresponding to the data frame, is the covariance matrix corresponding to the domain data frame cluster, The meaning of can refer to the above formula 1 and will not be elaborated herein.

[0108] S204. Use the target data frame as the data frame at the first moment.

[0109] In some embodiments, the server can retain the target data frame, that is, use the target data frame as the data frame of the target vehicle at the first moment, and delete other data frames in the first data frame except the target data frame.

[0110] Based on the above technical means, when there are data frames with the same acquisition time in the data of the target vehicle, the second data frame can be determined according to multiple data frames with the same acquisition time (i.e., the first data frame) and the adjacent data frames of the first data frame. The obtained second data frame can comprehensively reflect the running information of the target vehicle at multiple moments. Then, use the target data frame with the highest similarity to the second data frame in the first data frame as the data frame at the first moment, and delete other data frames in the first data frame except the target data frame. In this way, not only can the accuracy and reliability of the finally retained target data frame be ensured, but also the data duplication problem can be solved, and the overall quality of vehicle data can be improved.

[0111] In the embodiments of the present application, each data frame included in the first data frame may include different parameters. The following takes (1) the parameters included in multiple data frames satisfy the first judgment condition and (2) the parameters included in multiple data frames satisfy the second judgment condition as examples to introduce the data processing method provided by the present application.

[0112] (1) The parameters included in multiple data frames satisfy the first judgment condition.

[0113] In an alternative embodiment, the above S202 may include: when the parameters included in multiple data frames satisfy the first judgment condition, the server may determine a second data frame based on the adjacent data frames and the first data frame of the first data frame.

[0114] Wherein, each data frame may include continuous parameters and / or discrete parameters. A continuous parameter refers to a parameter that can take any value within a target interval, and a discrete parameter refers to a parameter whose values are finite or countable.

[0115] In the embodiments of the present application, the discrete parameters may include, but are not limited to, the positioning state of the vehicle, the vehicle state, and the charging state. The continuous parameters may include, but are not limited to, the drive motor speed, the drive motor torque, the drive motor temperature, the input voltage of the motor controller, the DC bus current of the motor controller, the longitude, the latitude, the vehicle speed, the cumulative mileage, the total voltage, the total current, and the state of charge (SOC).

[0116] On this basis, the first judgment condition includes at least one of the following:

[0117] 1-1. The range of at least one continuous parameter included in multiple data frames is greater than the range threshold.

[0118] In the embodiments of the present application, the range threshold corresponding to the continuous parameter may be determined according to actual needs, and no limitation is made thereto. For example, Table 1 shows the range thresholds of the continuous parameters included in the data frames.

[0119] Table 1 Range Thresholds of Continuous Parameters Included in Data Frames

[0120]

[0121] Exemplarily, in combination with Table 1 above, taking speed as an example. Assume that the speed corresponding to data frame 1 is 3 km / h, the speed corresponding to data frame 2 is 6 km / h, and the speed corresponding to the data frame is 4 km / h. The range of these three data frames is the difference (3 km / h) between the maximum value (6 km / h) and the minimum value (3 km / h). Since the range of the three data frames (3 km / h) is greater than the speed range threshold (2 km / h), it indicates that the first judgment condition is satisfied.

[0122] 1-2. The discrete parameters included in multiple data frames are different.

[0123] Exemplarily, taking the vehicle state as an example, the vehicle state may include a starting state, a shutting-down state, and other states. On this basis, if the vehicle state corresponding to data frame 1 is the starting state, the vehicle state corresponding to data frame 2 is the shutting-down state, and the vehicle state corresponding to data frame 3 is the starting state, it indicates that the first judgment condition is satisfied.

[0124] In some embodiments, when the parameters included in multiple data frames satisfy any one of the above first judgment conditions, the server may determine the second data frame with reference to the method in S202 above, which will not be elaborated here.

[0125] (2) The parameters included in multiple data frames satisfy the second judgment condition.

[0126] In an alternative embodiment, when the parameters included in multiple data frames satisfy the second judgment condition, the server may select any one of the data frames from the first data frame as the data frame at the first moment.

[0127] Among them, the second judgment condition includes that the range of each continuous parameter included in multiple data frames is less than or equal to the range threshold, and each discrete parameter included in multiple data frames is the same.

[0128] In some embodiments, when the parameters included in multiple data frames satisfy the second judgment condition, the server may randomly retain any one of the data frames and delete the other data frames in the first data frame except this data frame.

[0129] Exemplarily, taking the continuous parameter as the vehicle speed and the discrete parameters as the vehicle state and the charging state as an example. If the range of the vehicle speeds corresponding to data frame 1, data frame 2, and data frame 3 is less than the vehicle speed range threshold, and the vehicle states and charging states corresponding to data frame 1, data frame 2, and data frame 3 are completely the same, it indicates that the second judgment condition is satisfied. In this case, the server may retain data frame 1 and delete data frame 2 and data frame 3.

[0130] Based on the above technical means, the present application can determine different duplicate data processing strategies based on the preset conditions satisfied by the parameters included in multiple data frames. When the parameters in multiple data frames satisfy the first judgment condition, the operation of determining the second data frame is performed, avoiding the ineffective processing of data that does not meet specific requirements, and greatly improving the accuracy of data processing. When the parameters of multiple data frames satisfy the second judgment condition, any one is directly selected from the first data frame as the data frame at the first moment, without other complex processing, reducing the demand for hardware computing resources.

[0131] In an alternative embodiment, before performing the above S202, the data processing method provided by this application may further include: when the number of adjacent data frames of the first data frame is less than a threshold, the server may generate a third data frame based on the adjacent data frames of the first data frame. On this basis, the above S202 may be implemented as: the server may determine the second data frame based on the adjacent data frames of the first data frame, the first data frame, and the third data frame.

[0132] In some embodiments, when the number of adjacent data frames of the first data frame is less than a threshold, the server may use a preset oversampling method to generate a third data frame based on the adjacent data frames of the first data frame. Among them, the third data frame includes multiple new data frames generated by the preset oversampling method. After that, the server may refer to the method in the above S202 to determine the second data frame based on the first data frame, the adjacent data frames of the first data frame, the third data frame, and the quantities of the first data frame, the adjacent data frames of the first data frame, and the third data frame, which will not be elaborated here.

[0133] In the embodiments of this application, the preset oversampling method may include random linear interpolation with noise addition oversampling, random oversampling, generative adversarial network oversampling, etc., which are not limited herein.

[0134] In some other embodiments, the server may add the first data frame and the adjacent data frames of the first data frame to the domain data frame cluster and calculate the covariance matrix corresponding to the domain data frame cluster. After that, the server may calculate the determinant of the covariance matrix, and when the absolute value of the determinant is less than the determinant threshold, use the preset oversampling method to generate a third data frame based on the adjacent data frames of the first data frame.

[0135] Exemplarily, taking random linear interpolation with noise addition oversampling as an example. The server may successively use random linear interpolation with noise addition oversampling to generate a new data frame and add it to the domain data frame cluster until the absolute value of the determinant of the covariance matrix corresponding to the domain data frame cluster is greater than the determinant threshold. Among them, the new data frame may be determined by the following formula 3.

[0136] (Formula 3).

[0137] Among them, is the new data frame generated by oversampling, is the random linear interpolation weight, ; is a randomly selected data frame from the adjacent data frames of the first data frame and is not the last data frame, is the data frame is the adjacent data frame of; is a random vector, and Each component in follows a uniform distribution and is randomly generated on is the maximum absolute value vector of noise. It should be noted that the above random variables, such as , etc. are randomly regenerated in each round of oversampling.

[0138] The following will be combined with Figure 3 , taking random linear interpolation with noise addition and oversampling as an example, to introduce the method in the above embodiments. As Figure 3 shown, the method includes: S301 - S308.

[0139] S301. Determine that the acquisition time of the first data frame is t, and the time domain radius threshold is .

[0140] S302. Add the data frames whose acquisition times are within the range of to the domain data frame cluster, and determine the data frames in the domain data frame cluster except the first data frame as the centered - removed domain data frame cluster.

[0141] S303. Calculate the absolute value DA of the determinant of the covariance matrix corresponding to the domain data frame cluster.

[0142] S304. Judge whether DA is less than . If so, execute S305; if not, execute S306.

[0143] S305. Perform random linear interpolation with noise addition and oversampling within the centered - removed domain data frame cluster, generate new data frames and add them to the domain data frame cluster, and then execute S303.

[0144] S306. Determine the second data frame based on each data frame in the domain data frame cluster.

[0145] S307. Calculate the Mahalanobis distance between each data frame in the first data frame and the second data frame.

[0146] S308. Retain the data frame in the first data frame with the smallest Mahalanobis distance from the second data frame, and delete the other data frames in the first data frame.

[0147] Based on the above technical solution, when the number of adjacent data frames of the first data frame is insufficient, the present application can combine the adjacent data frames of the first data frame to generate a third data frame. Alternatively, when the absolute value of the determinant of the covariance matrix corresponding to the first data frame and the adjacent data frames of the first data frame is less than the absolute value threshold, a third data frame is generated to avoid the determinant of the covariance matrix being too small to calculate the Mahalanobis distance between each data frame in the first data frame and the second data frame. In addition, during the subsequent determination of the second data frame, the third data frame is used as supplementary information and combined with the adjacent data frames of the first data frame and the first data frame to further optimize the determination process of the second data frame, enabling the second data frame to integrate more dimensional information and improving the overall efficiency of data utilization.

[0148] In an alternative embodiment, before executing S201 above, the data processing method provided by the present application may further include: The server can obtain multiple data frames collected by the target vehicle at different times. Then, for each data frame among the multiple data frames collected by the target vehicle at different times, the server can determine the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame. Finally, for any two adjacent data frames among the multiple data frames collected by the target vehicle at different times, the server can determine the data frames belonging to the same driving segment based on the vehicle running states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames. On this basis, S201 above may specifically include: The server can obtain the first data frame from among the multiple data frames included in the same driving segment.

[0149] In the embodiments of the present application, the vehicle state refers to the state of the vehicle in different working modes, and the vehicle running state includes different states during the vehicle's operation.

[0150] In some embodiments, the vehicle running state may include a first running state and a second running state. The vehicle state may include a start state and an unstarted state. The charging state may include a driving charging state and an uncharged state, and the driving charging state refers to the state when the battery recovers energy through the energy recovery system during vehicle driving.

[0151] On the above basis, after the server obtains multiple data frames collected by the target vehicle at different times, for each data frame among the multiple data frames collected by the target vehicle at different times, the server can determine the vehicle running state corresponding to the data frame with the vehicle state being the start state and the charging state being the driving charging state or the uncharged state as the first running state, and determine the vehicle running state corresponding to the data frame with the vehicle state being the unstarted state as the second running state.

[0152] After that, for any two adjacent data frames among multiple data frames collected for the target vehicle at different times, if the vehicle operating states corresponding to any two adjacent data frames are different, then the two adjacent data frames do not belong to the same driving segment. If the vehicle operating states corresponding to any two adjacent data frames are the same but the time interval between any two adjacent data frames is greater than the time interval threshold, then the two adjacent data frames do not belong to the same driving segment. After that, for the data frames in each driving segment, the server can delete the data frames in the second operating state. Finally, the server can obtain the first data frame from the multiple data frames included in the same driving segment with reference to the method in S201 above, which will not be elaborated here.

[0153] In the embodiments of the present application, the time interval threshold can be 500 seconds, 600 seconds, etc., which is not limited herein.

[0154] In one example, the data frames in the first operating state are defined as driving frames, and the data frames in the second operating state are defined as non-driving frames. On this basis, the method for determining the data frames belonging to the same driving segment will be introduced in detail below in combination with the above various embodiments, as Figure 4 shown, the method includes: S401 - S412.

[0155] S401. Sort the multiple data frames of the target vehicle according to the acquisition time.

[0156] Specifically, the server can arrange the multiple data frames of the target vehicle in ascending order of the acquisition time.

[0157] S402. Determine whether the initial data frame (i.e., the 0th data frame) of the target vehicle is a driving frame. If so, execute S403; if not, execute S404.

[0158] S403. Determine the initial data frame as the starting frame of the new driving segment.

[0159] S404. Initialize the frame index i = 1.

[0160] S405. Determine whether the frame index i is less than or equal to i_max. If so, execute S406; if not, execute S411.

[0161] Wherein, i_max is the number of multiple data frames.

[0162] S406. Determine whether the i-th data frame is a driving frame. If so, execute S407; if not, execute S410.

[0163] S407. Determine whether the (i - 1)-th data frame is a driving frame. If so, execute S408; if not, execute S409.

[0164] S408. Determine whether the time interval between the i-th data frame and the (i - 1)-th data frame is greater than the time interval threshold. If so, execute S409; if not, execute S410.

[0165] S409. Determine the i-th data frame as the starting frame of a new driving segment.

[0166] S410. Let the frame index i = i + 1.

[0167] S411. Determine the i_max-th data frame as the starting frame of a new segment.

[0168] S412. Determine the data frames between two adjacent starting frames of new segments, and the data frame with a lower order among the two adjacent starting frames of new segments as belonging to the same driving segment.

[0169] Exemplarily, if data frame 1 is the starting frame of a new segment and data frame 4 is the starting frame of a new segment, then data frames 1, 2, and 3 belong to the same driving segment. After that, the server can also determine the driving segment numbers of each data frame. For example, the driving segment numbers of data frames 1, 2, and 3 are 1, and the driving segment number of data frame 4 is 2.

[0170] Based on the above technical means, the present application can determine the vehicle operation state by comprehensively considering the vehicle state and charging state corresponding to each data frame. This method can more comprehensively and accurately reflect the actual operation of the vehicle at different times. Then, based on the vehicle operation state and time interval corresponding to any two adjacent data frames, it is determined whether they belong to the same driving segment. This method can reasonably divide the driving process of the vehicle, and then obtain data with clear travel boundaries. At the same time, accurate division of driving segments is of great significance for analyzing the driving behavior, energy consumption, driving route, etc. of the vehicle. At the same time, by deleting the data frames in the second operation state, the storage space can be saved and the efficiency of subsequent data processing can be improved. In addition, by obtaining the first data frame from multiple data frames included in the same driving segment, such an operation method makes the data processing more targeted and can improve the data processing efficiency.

[0171] In an alternative embodiment, before determining the vehicle operation state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame, the method provided by the present application may further include: for each data frame included in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meet the preset correction conditions, correct the vehicle state and / or charging state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame.

[0172] Among them, the correction conditions include any one of the following:

[0173] 2-1. The vehicle state corresponding to each data frame included in the same driving segment is other vehicle states.

[0174] Among them, other vehicle states refer to vehicle states other than the starting state and the shutting-down state.

[0175] 2-2. The vehicle state corresponding to each data frame included in the same driving segment is the starting state and the charging state is other charging states. Among them, other charging states refer to charging states other than the driving charging state and the non-charging state.

[0176] It should be noted that due to equipment failures and other reasons in the vehicle, there may be missing values, outliers, or meaningless values in the vehicle state and charging state corresponding to each data frame. Therefore, it is necessary to correct the vehicle state and charging state corresponding to each data frame.

[0177] Combining the above content, the following takes (1) correcting the vehicle state corresponding to the data frame and (2) correcting the charging state corresponding to the data frame as examples for introduction.

[0178] (1) Correct the vehicle state corresponding to the data frame.

[0179] In some embodiments, for each data frame included in the same driving segment, when the vehicle state corresponding to each data frame is other vehicle states, the server can correct the vehicle state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame. Among them, the following explanations are made for the correction of the vehicle state:

[0180] 1.1. For safety considerations, parking charging is usually carried out when the vehicle is shut down. Therefore, it is default that when the charging state is parking charging, the vehicle state is the shutting-down state.

[0181] 1.2. When the charging state is the driving charging state, it indicates that the vehicle is in regenerative braking, and the vehicle state should be the starting state.

[0182] 1.3. When the total current is positive, it indicates that the power battery is outputting electrical energy, and the vehicle state should be the starting state.

[0183] Combining the above explanations, taking the correction of the vehicle state corresponding to each data frame included in the target driving segment as an example, as Figure 5 shown, the method includes: S501 - S513.

[0184] S501. Determine whether there is a data frame with a charging state of parking charging state in the target driving segment. If so, execute S502; if not, execute S503.

[0185] S502. Determine the vehicle state corresponding to each data frame included in the target driving segment as the off state.

[0186] S503. Determine whether there is a data frame with a charging state of driving charging state in the target driving segment. If so, execute S504; if not, execute S505.

[0187] S504. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0188] S505. Determine whether more than half of the data frames in the target driving segment have a vehicle speed greater than 1 km / h. If so, execute S506; if not, execute S507.

[0189] S506. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0190] S507. Determine whether more than half of the data frames in the target driving segment have a vehicle speed less than or equal to 1 km / h. If so, execute S508; if not, execute S511.

[0191] S508. Determine whether more than half of the data frames in the target driving segment have a total current greater than 1 A. If so, execute S509; if not, execute S510.

[0192] S509. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0193] S510. Determine the vehicle state corresponding to each data frame included in the target driving segment as the off state.

[0194] S511. Determine whether more than half of the data frames in the target driving segment have a total current greater than 1 A. If so, execute S512; if not, execute S513.

[0195] S512. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0196] S513. Determine the vehicle state corresponding to each data frame included in the target driving segment as the invalid state.

[0197] In the embodiments of the present application, the invalid state can be represented by 255, or can be represented by other means, which is not limited herein.

[0198] (2) Correct the charging state corresponding to the data frame.

[0199] In some embodiments, for each data frame included in the same driving segment, when the vehicle state corresponding to each data frame is the starting state and the charging state is other charging states, the server may correct the charging state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame. Among them, the following explanations are made for the correction of the charging state:

[0200] 2.1. When the vehicle is in the starting state (for example, the vehicle speed corresponding to more than half of the data frames in the target driving segment is greater than 1 km / h), uniformly correct the charging state to the non-charging state.

[0201] 2.2. When the information obtained from the charging state, vehicle speed, and total current is contradictory or all are missing values, correct the charging state to 255, that is, the invalid state.

[0202] 2.3. If it is determined according to the vehicle speed, total current, and charging state that the vehicle is charging while parked in the starting state, respect the data record and allow the contradiction between the vehicle state and the charging state, and correct the charging state to the parked charging state.

[0203] 2.4. The charging completed state may occur both after the parked charging is full and at the initial stage of vehicle startup, and cannot be used for judging the driving state. When this value exists and the charging state cannot be determined through other parameters, the charging state remains the charging completed state.

[0204] Combined with the above explanations, taking the correction of the charging state corresponding to each data frame included in the target driving segment as an example, as Figure 6 shown, the method includes: S601 - S613.

[0205] S601. Judge whether there are more than half of the data frames in the target driving segment whose corresponding vehicle speeds are greater than 1 km / h. If so, execute S602; if not, execute S603.

[0206] S602. Determine the charging state corresponding to each data frame included in the target driving segment as the non-charging state.

[0207] S603. Judge whether there are more than half of the data frames in the target driving segment whose corresponding total currents are less than or equal to -1 A. If so, execute S604; if not, execute S607.

[0208] S604. Judge whether there is a data frame with a charging state of parked charging state in the target driving segment. If so, execute S605; if not, execute S606.

[0209] S605. Determine the charging state corresponding to each data frame included in the target driving segment as the parked charging state.

[0210] S606. Determine the charging status corresponding to each data frame included in the target driving segment as an invalid charging status.

[0211] S607. Determine whether there are more than half of the data frames in the target driving segment whose total current is greater than -1A. If so, execute S608; if not, execute S611.

[0212] S608. Determine whether there is a data frame with a charging status of fully charged in the target driving segment. If so, execute S609; if not, execute S610.

[0213] S609. Determine the charging status corresponding to each data frame included in the target driving segment as a fully charged status.

[0214] S610. Determine the charging status corresponding to each data frame included in the target driving segment as an uncharged status.

[0215] S611. Determine whether there is a data frame with a charging status of parked charging in the target driving segment. If so, execute S612; if not, execute S613.

[0216] S612. Determine the charging status corresponding to each data frame included in the target driving segment as a parked charging status.

[0217] S613. Determine the charging status corresponding to each data frame included in the target driving segment as an invalid charging status.

[0218] Based on the above technical means, the present application can correct problems such as value gaps, anomalies, and inaccuracies in the charging status and vehicle status caused by various factors such as sensor errors and signal interference, improve data quality, and provide support for subsequent determination of the vehicle operating status corresponding to each data frame.

[0219] In an optional implementation manner, the data processing method provided by the present application may further include: for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, the server can successively use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained.

[0220] Among them, the missing value includes a parameter with a null value or a parameter with a value exceeding a preset range.

[0221] In the embodiments of the present application, each parameter included in the data frame corresponds to a threshold range (i.e., a reasonable range), and the threshold range is determined based on the specification parameters and common sense of the components in the vehicle to which the data frame belongs. For example, Table 2 shows the threshold ranges of each parameter. Referring to Table 2, for each parameter, the server can replace the data with a parameter value exceeding the corresponding threshold range with a missing value.

[0222] Table 2 Threshold ranges of each parameter

[0223]

[0224] Among them, the order of the sub-imputation models for imputing missing values is related to the accuracy index of the sub-imputation models, and the accuracy index is used to characterize the accuracy rate of the imputation results obtained by using the sub-imputation models.

[0225] The embodiments of the present application do not limit the sub-imputation models. For example, the sub-imputation models may include a mean imputation model, a random hot deck method, a regression imputation model, a time series imputation model, etc., and no limitation is imposed thereon. Among them, the regression imputation model may include a LightGBM model, a K-nearest neighbor model, a random forest, etc. The time series imputation model may include an adjacent frame linear interpolation model, a cubic Lagrange interpolation model, etc.

[0226] In the embodiments of the present application, the accuracy index may include the root mean squared error (RMSE), the mean squared error (MSE), etc., and no limitation is imposed thereon.

[0227] In some embodiments, multiple integrated imputation models are deployed in the server, each integrated imputation model includes multiple sub-imputation models, and each parameter included in the data frame corresponds to an integrated imputation model. On this basis, taking the target parameter included in the target data frame in the first data frame as a missing value as an example, the method for performing imputation processing by using the target integrated imputation model corresponding to the target parameter is introduced as follows. As Figure 7 shown, the method includes: S701 - S707.

[0228] S701. Sort each sub-imputation model according to the accuracy index of each sub-imputation model in the target integrated imputation model.

[0229] Among them, the larger the accuracy index of the sub-imputation model, the more forward the order of the sub-imputation model, that is, the sub-imputation model with a high accuracy index can be preferentially used for imputation processing.

[0230] S702. Initialize the sub-imputation model index i = 1.

[0231] S703. Determine whether the sub-imputation model index i is less than or equal to the total number i_max of sub-imputation models in the target integrated imputation model. If so, execute S704; if not, execute S707.

[0232] S704. Determine whether other parameters in the target data frame except the target parameter meet the input parameters required by the i-th sub-imputation model. If so, execute S705; if not, execute S706.

[0233] S705. Use the i-th sub-imputation model to perform imputation processing on the target parameter.

[0234] S706. Let the sub-imputation model index i = i + 1, and execute S703.

[0235] S707. Use other imputation models except the sub-imputation models in the integrated imputation model to perform imputation processing on the target parameter.

[0236] In the embodiments of the present application, other imputation models may include a wide-range linear interpolation model, a nearest neighbor imputation model, etc., and are not limited thereto.

[0237] In one example, taking the longitude and latitude in the target data frame as missing values, the server can use the wide-range linear interpolation model to perform imputation processing on the longitude and latitude to obtain an imputation result. If the imputation using the wide-range linear interpolation model fails, the server can use the nearest neighbor imputation model to perform imputation processing on the longitude and latitude.

[0238] In another example, taking the cumulative mileage in the target data frame as a missing value, the server can determine the imputation value of the cumulative mileage corresponding to the target data frame based on the cumulative mileage corresponding to the previous data frame adjacent to the target data frame, the target data frame, the vehicle speed corresponding to the previous data frame adjacent to the target data frame, and the acquisition time corresponding to the previous data frame adjacent to the target data frame. For example, the imputation value of the cumulative mileage corresponding to the target data frame satisfies the following formula 4.

[0239] (Formula 4).

[0240] Where is the imputation value of the cumulative mileage corresponding to the target data frame, is the cumulative mileage corresponding to the previous data frame adjacent to the target data frame, is the acquisition time of the target data frame, is the acquisition time of the previous data frame adjacent to the target data frame, is the vehicle speed corresponding to the target data frame, is the vehicle speed corresponding to the previous data frame adjacent to the target data frame. If is greater than a predetermined threshold (such as 1 / 240 h), or is a missing value, the nearest neighbor imputation value can be used to perform imputation processing on the cumulative mileage corresponding to the target data frame.

[0241] Based on the above technical means, the sub-imputation model with high precision indicators can be preferentially used to impute missing values. When there are missing values in the input parameters and the current sub-imputation model cannot be used, the next sub-imputation model in the integrated imputation model is used sequentially until the imputation result is obtained. In this way, high-precision missing value imputation can be achieved as much as possible, and the feasibility of imputation can be guaranteed. The advantages and disadvantages of different sub-imputation models in terms of precision and usability can be complementary. After all sub-models have attempted to impute the missing values to be imputed, there may still be some missing values that all sub-imputation models fail to impute. In this case, a backup imputation model (i.e., other imputation models) can be used for imputation.

[0242] The above introduced the use of sub-imputation models in the integrated imputation model to process missing values. The following introduces how to construct an integrated imputation model.

[0243] In the embodiments of this application, take the sub-imputation model including a regression imputation model and a time series imputation model as an example. The regression imputation model includes the LightGBM model, the K-nearest neighbor model, and the random forest. The time series imputation model includes the adjacent frame linear interpolation model and the cubic Lagrange interpolation model.

[0244] It should be noted that the time series imputation model can be directly used for missing value imputation, while the regression imputation model needs to be trained before it can be used. Therefore, the model training mentioned below is mainly for the regression imputation model, and the time series imputation model can directly calculate the imputed value corresponding to the missing value based on a preset formula. For example, the adjacent frame linear interpolation model satisfies the following formula 5, and the cubic Lagrange interpolation model satisfies the following formula 6.

[0245] (Formula 5).

[0246] Among them, is the acquisition time of the data frame where the missing value is located, is the acquisition time of the previous data frame of the data frame where the missing value is located, is the acquisition time of the next data frame of the data frame where the missing value is located, ; is the value of the parameter corresponding to the missing value in the previous data frame, is the value of the parameter corresponding to the missing value in the next data frame, is the imputed value corresponding to the missing value.

[0247] (Formula 6).

[0248] Among them, is the acquisition time of the th data frame before the data frame where the missing value is located, is the The acquisition time of a data frame is the value of the parameter corresponding to the missing value in the th data frame.

[0249] In some embodiments, each data frame contains multiple parameters, and each type of parameter corresponds to an integrated imputation model. Taking the construction of the target integrated imputation model corresponding to the target parameter as an example, as Figure 8 shown, the target integrated imputation model can be obtained through the following S801 - S806.

[0250] S801. Determine the training data frames and test data frames from multiple data frames collected at different times.

[0251] Among them, the training data frames and test data frames include multiple data frames in which the target parameter is not a missing value. The training set is used to train the sub - imputation model, and the test set is used to determine the accuracy index of the trained sub - imputation model.

[0252] Specifically, the server can use the random sampling without replacement or random sampling with replacement method to obtain the sampled data frames from multiple data frames collected at different times. Then, the server can divide the sampled data frames into training data frames and test data frames based on a preset division ratio. Then, for each data frame in the test data frames, the server can obtain the acquisition time of the data frame, and the acquisition times of the two data frames before and after the data frame in multiple data frames.

[0253] The embodiments of the present application do not limit the preset division ratio. For example, the sampled data frames can be divided according to a ratio of 5:1, or can be divided according to a ratio of 4:1.

[0254] In the embodiments of the present application, it is possible to determine whether to use the random sampling without replacement or the random sampling with replacement method based on the number of multiple data frames and the number of vehicles to which the multiple data frames belong. For example, if the ratio of the number of multiple data frames to the number of vehicles is greater than or equal to the quantity threshold, the random sampling without replacement is used; if the ratio of the number of multiple data frames to the number of vehicles is less than the quantity threshold, the random sampling with replacement is used.

[0255] In the embodiments of the present application, before training the sub - imputation model, it is also necessary to perform standardization processing on other parameters in the training data frames except the target parameter. Taking the K - nearest neighbor model as an example. Before training the K - nearest neighbor model corresponding to the target parameter, the server can first perform standardization processing on other parameters except the target parameter, and the standardization calculation formula can refer to the following formula 7.

[0256] (Formula 7).

[0257] Among them, is a parameter among multiple parameters, is the corresponding parameter mean in the training data frame, is the corresponding parameter standard deviation in the training data frame, is the standardized parameter.

[0258] S802. Screen out the optimal feature parameters from other parameters included in the training data frame except the target parameter, and use them as the input features of the sub-imputation model.

[0259] Specifically, the server can use a preset feature selection method to screen out the optimal features from each parameter included in the training data frame as the input features of the sub-imputation model, so as to delete redundant features and ensure the accuracy and efficiency of the sub-imputation model.

[0260] In the embodiments of the present application, the preset feature selection method may include methods such as manual feature selection and backward sequential feature elimination, which are not limited herein. For example, the method of using the backward sequential feature elimination method to determine the optimal features can refer to the following embodiments as Figure 9 shown, which will not be elaborated herein.

[0261] S803. For each original sub-imputation model, use the optimal feature parameters in the training data frame to train the original sub-imputation model to obtain a trained sub-imputation model.

[0262] Specifically, the server can input the optimal feature parameters in the training data frame into each original imputation model to obtain each trained sub-imputation model.

[0263] S804. For each trained sub-imputation model, input the optimal feature parameters in the test data frame into the trained sub-imputation model to obtain the imputation result corresponding to the target parameter.

[0264] S805. For each trained sub-imputation model, based on the target parameter included in the test data frame and the imputation result corresponding to the target parameter, determine the accuracy index of the trained sub-imputation model.

[0265] Exemplarily, taking the accuracy index as the root mean square error RMSE, the smaller the RMSE, the higher the model accuracy. Table 3 shows the root mean square error RMSE of different sub-imputation models under different parameters.

[0266] Table 3 Root Mean Square Error RMSE of Different Sub-imputation Models under Different Parameters

[0267]

[0268] S806. Combine in the order of the accuracy indexes of each trained sub-imputation model to obtain the target integrated imputation model corresponding to the target parameter.

[0269] Exemplarily, taking the driving motor speed as the target parameter, combined with the above Table 3. Since the smaller the RMSE, the higher the model accuracy, therefore, as Figure 10 shown, the order of the sub-interpolation models included in the target integrated interpolation model is Random Forest, K-Nearest Neighbor model, LightGBM model, adjacent frame linear interpolation model, cubic Lagrange interpolation model. In addition, the wide-range linear interpolation model and the nearest neighbor interpolation model can also be used as backup models and integrated into the target integrated interpolation model.

[0270] As Figure 10 shown, in the case where the target parameter in data frame A is a missing value, the server can, in the order of the integrated interpolation model, successively use the Random Forest, K-Nearest Neighbor model, LightGBM model, adjacent frame linear interpolation model, cubic Lagrange interpolation model, wide-range linear interpolation model, and nearest neighbor interpolation model to interpolate the missing value until the interpolation result corresponding to the target parameter is obtained. Among them, the parameters input to each sub-interpolation model include other parameters in data frame A except the target parameter, and the target parameters included in the previous data frame B and the next data frame C adjacent to data frame A.

[0271] Based on the above technical means, the present application can solve the situation where the parameters included in the data frame have missing values by constructing an integrated interpolation model. Compared with directly deleting the data frame containing the missing value, using the interpolation model for processing can retain more valid data and reduce data waste.

[0272] In some embodiments, taking the backward sequential feature elimination method as an example, as Figure 9 shown, using the backward sequential feature elimination method to obtain the optimal features may specifically include: Step 1 - Step 10.

[0273] Step 1: Obtain a sub-training data frame for optimal feature selection from the training data frame.

[0274] It should be noted that based on the feature selection method of backward sequential feature elimination, the computational complexity is extremely large. Therefore, the server can randomly sample from the training set to obtain a sub-training data frame for optimal feature selection. Among them, the sample size of the sub-training data frame can be determined according to the actual computing resources and is not limited thereto.

[0275] Step 2: Use the sub-training data frame to train a sub-interpolation model and determine the initial accuracy index of the sub-interpolation model.

[0276] In the embodiments of the present application, the accuracy metrics may include root mean squared error (RMSE), mean squared error (MSE), etc., which are not limited herein.

[0277] Specifically, the server may input other parameters in the sub-training data frame except the target parameter into the sub-imputation model to obtain the target parameter predicted by the model. Then, the server may calculate the accuracy metric of the sub-imputation model based on the target parameter in the sub-training data frame and the target parameter predicted by the model.

[0278] Step 3: Optimize the hyperparameters of the sub-imputation model by using a preset hyperparameter optimization algorithm.

[0279] Among them, hyperparameters refer to the parameters set before training the model, which are used to control the behavior and performance of the model. The selection of hyperparameters can affect the training speed, convergence, capacity, and generalization ability of the model, etc.

[0280] Specifically, the server may first determine the target hyperparameters and the search range corresponding to each target hyperparameter. Then, the server uses a preset hyperparameter optimization algorithm to determine the optimal value of each hyperparameter from the search range corresponding to each target hyperparameter. Table 4 shows the target hyperparameters to be optimized for the K-nearest neighbor model, random forest, and LightGBM model.

[0281] Among them, the target hyperparameters corresponding to the K-nearest neighbor model may include the number of neighboring data (n_neighbors), weights, and distance metric parameter (p). The target hyperparameters corresponding to the random forest may include the number of iterations (n_estimators), the maximum depth of the tree (max_depth), the minimum number of samples in the leaf node (min_samples_leaf), the maximum number of features (max_features), and the maximum number of samples in each tree (max_samples). The target hyperparameters corresponding to the LightGBM model may include n_estimators, the number of leaf nodes in each tree (num_leaves), max_depth, learning rate, the minimum number of samples required for the leaf node (min_child_samples), the sample ratio used when training each tree (subsample), and the subsample frequency, that is, how many rounds of iteration to perform a subsampling (subsample_freq).

[0282] Table 4 Target hyperparameters to be optimized for the K-nearest neighbor model, random forest, and LightGBM model

[0283]

[0284] In the embodiments of the present application, the preset hyperparameter optimization algorithm may include a random search algorithm, a Bayesian optimization algorithm, a grid search algorithm, etc., and no limitation is imposed thereon.

[0285] Exemplarily, taking the sub-imputation model as the LightGBM model and the preset hyperparameter optimization algorithm as the Bayesian optimization algorithm as an example. The server can use the Python language to call the lightgbm package to implement the LightGBM model. Then the server can call the Bayesian optimization algorithm (tree-structured parzen estimator, TPE) provided by the optuna package to search for the optimal hyperparameters of the LightGBM model. Among them, in the process of searching for hyperparameters, the accuracy index of each group of hyperparameters can be obtained by using the K-fold cross-validation method, and the accuracy index can be the root mean square error RMSE.

[0286] Step 4: Initialize the parameter sequence i = 1.

[0287] Step 5: Determine whether i is greater than the remaining parameters i _ max in the sub-training data frame. If so, execute Step 8; if not, execute Step 6.

[0288] Step 6: Use the other parameters in the sub-training data frame except the i-th parameter to train the sub-imputation model, and determine the accuracy index corresponding to the sub-imputation model.

[0289] Step 7: Let i = i + 1.

[0290] Step 8: Determine the redundant parameters so that the sub-imputation model trained based on the other parameters except the redundant parameters has the highest accuracy index, and delete the redundant parameters in the sub-training data frame.

[0291] Step 9: Determine the remaining parameter i _ max = 1 or whether the continuous decline times of the accuracy index of the sub-imputation model are greater than the times threshold. If so, execute Step 10; if not, execute Step 3.

[0292] Among them, the times threshold can be 5 times, 6 times, etc., and no limitation is imposed thereon in the present application.

[0293] Step 10: Use the input features corresponding to the sub-imputation model with the highest accuracy index as the optimal parameter features of the sub-training data frame.

[0294] Combining the above steps 1 - 10, it can be understood that backward sequence feature elimination can be iterated multiple times. By deleting one parameter in the sub-training in each iteration, the original parameters in the sub-training data frame are refined into optimal parameter features to improve the accuracy of the sub-imputation model. The finally obtained optimal parameter features are the input features corresponding to the model with the highest accuracy index in all iterations.

[0295] In the embodiments of the present application, the optimal feature parameters (i.e., the model input features) may vary depending on the training data frame and the regression imputation model. Table 5 shows the feature selection results of the regression imputation sub-model. In Table 5, the driving motor speed - 1 is the driving motor speed of the data frame before the target data frame, the driving motor speed + 1 is the driving motor speed of the data frame after the target data frame, and so on for other cases, which will not be elaborated here. The input feature is the optimal feature, and the target parameter is the output parameter of the model, that is, the target parameter predicted by the model can be obtained after inputting the input feature into the model.

[0296] Table 5 Feature Selection Results of the Regression Imputation Sub-Model

[0297]

[0298] Continued Table 5

[0299]

[0300] Based on the above technical solution, the present application can delete redundant feature parameters and use the optimal feature parameters as input features to train the model, thereby improving the processing efficiency and output accuracy of the sub-imputation model.

[0301] In an alternative embodiment, before performing the above S201, the data processing method provided by the present application further includes: the server can decode each original data frame according to a preset decoding rule to obtain multiple decoded data frames, and each decoded data frame corresponds to a vehicle identification code. Then, the server can store the data frames with the same vehicle identification code correspondingly to obtain the data frames corresponding to each vehicle. On this basis, the above S201 may include: for the target vehicle among multiple vehicles, the server can obtain the first data frame from the data frames corresponding to the target vehicle.

[0302] In some embodiments, the server stores the original data frames collected by different vehicles at different times, and each original data frame is an encoded data frame. Based on this, the server can decode each original data frame according to a preset decoding rule. After that, the server can segment the decoded data frames according to the vehicle identification code (VID) in each decoded data frame to obtain the data frames corresponding to each vehicle. Then, for each vehicle, the server can store the data frames corresponding to the vehicle in a structured manner according to the collection time of the data frames corresponding to the vehicle.

[0303] Exemplarily, Table 6 shows the encoding rules for each parameter in the original data frame based on the GB / T 32960 standard. The server can decode the parameters in each data frame (such as scaling, offset) based on the encoding rules to obtain the true values of each parameter in common units. Combining Table 6 below, taking the driving motor speed as an example, the effective value range of the encoded driving motor speed is between 0 and 65531, and its data offset is 20000. The server can subtract the data offset of 20000 from the encoded driving motor speed to obtain the decoded driving motor speed, whose range is between -20000 r / min and 45531 r / min. After that, the server can store the decoded data in the structured data table shown in Table 7. Each row in the structured data table is a data frame, and each data frame contains the various parameters of the vehicle at a certain collection time.

[0304] Table 6 Encoding rules for each parameter in the original data frame based on the GB / T 32960 standard

[0305]

[0306] Continued Table 6

[0307]

[0308] Table 7 Structured data table

[0309]

[0310] Based on the above technical solutions, by decoding the original data frames according to the preset decoding rules and storing the data frames corresponding to the vehicle identification codes, the effective classification of vehicle data is realized, which is convenient for subsequent data management and analysis of specific vehicles. For the target vehicle among multiple vehicles, the first data frame can be obtained from its corresponding data frames. This method makes the processing of specific vehicle data more efficient and accurate, and improves the data processing efficiency.

[0311] The above has introduced the processing methods for data frames (and duplicate data) with the same acquisition time, the processing methods for missing values in the parameters included in the data frames, the driving segment division method, and the method for obtaining the data frames corresponding to each vehicle. The following will introduce the data processing method provided by this application in combination with the above various embodiments. As Figure 11 shown, the data processing method provided by this application mainly includes data structuring processing (S1101 - S1102), driving segment division (S1103 - S1104), outlier processing (S1105 - S1106)), and missing value imputation (S1107 - S1108). Specifically, this method can include:

[0312] S1101. Decode the original data according to a preset decoding rule.

[0313] S1102. Integrate the decoded data into a structured data table with the acquisition time series as rows and each parameter as columns. Each row in the structured data table is a data frame.

[0314] S1103. Correct the vehicle status and charging status corresponding to each data frame, and determine the vehicle running status corresponding to each data frame based on the vehicle status and charging status corresponding to each data frame.

[0315] S1104. For any two adjacent data frames among multiple data frames, determine whether the two adjacent data frames belong to the same driving segment based on the vehicle running status corresponding to the two adjacent data frames and the time interval between the two adjacent data frames.

[0316] S1105. For each data frame in the same driving segment, determine the parameter that exceeds the threshold range in the data frame as an outlier, and replace the parameter with a missing value.

[0317] S1106. Obtain the first data frame from the same driving segment, and determine the second data frame based on the adjacent data frame and the first data frame of the first data frame, and retain the data frame with the highest similarity to the second data frame in the first data frame.

[0318] S1107. Construct an integrated imputation model.

[0319] S1108. For each data frame in the same driving segment, when the parameter included in the data frame is a missing value, use the integrated imputation model to process the missing value.

[0320] Figure 12 is a schematic structural diagram of a data processing device provided by an embodiment of this application. As Figure 12As shown in the figure, the device includes: an acquisition unit 1201, a first determination unit 1202, a second determination unit 1203, and a processing unit 1204.

[0321] The acquisition unit 1201 is configured to acquire a first data frame, and the first data frame includes a plurality of data frames collected by the target vehicle at a first moment.

[0322] The first determination unit 1202 is configured to determine a second data frame based on the adjacent data frame of the first data frame and the first data frame. The adjacent data frame of the first data frame includes data frames collected by the target vehicle at adjacent moments of the first moment.

[0323] The second determination unit 1203 is configured to determine a target data frame with the highest similarity to the second data frame from the first data frame.

[0324] The processing unit 1204 is configured to use the target data frame as the data frame at the first moment.

[0325] In a possible way, the first determination unit 1202 is further configured to determine a second data frame based on the adjacent data frame of the first data frame and the first data frame when the parameters included in the plurality of data frames meet the first judgment condition.

[0326] In a possible way, the first determination unit 1202 is further configured to generate a third data frame based on the adjacent data frame of the first data frame when the number of adjacent data frames of the first data frame is less than a threshold. Then, a second data frame is determined based on the adjacent data frame of the first data frame, the first data frame, and the third data frame.

[0327] In a possible way, the second determination unit 1203 is further configured to, for each data frame in the first data frame, determine the degree of difference between the data frame and the second data frame, and use the data frame with the smallest degree of difference between the first data frame and the second data frame as the target data frame.

[0328] In a possible way, the acquisition unit 1201 is further configured to acquire a plurality of data frames collected by the target vehicle at different moments.

[0329] In a possible way, the first determination unit 1202 is further configured to, for each data frame in the plurality of data frames collected by the target vehicle at different moments, determine the vehicle running state corresponding to the data frame based on the vehicle state and the charging state corresponding to the data frame.

[0330] In a possible way, the first determination unit 1202 is further configured to determine, for any two adjacent data frames among a plurality of data frames collected by the target vehicle at different times, whether the two adjacent data frames belong to the same driving segment based on the vehicle running states corresponding to the two adjacent data frames and the time interval between the two adjacent data frames.

[0331] In a possible way, the obtaining unit 1201 is further configured to obtain a first data frame from among the plurality of data frames included in the same driving segment.

[0332] In a possible way, the processing unit 1204 is further configured to, for each data frame included in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meets a preset correction condition, correct the vehicle state and / or charging state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame.

[0333] In a possible way, the processing unit 1204 is further configured to, for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, sequentially use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained.

[0334] Figure 13 This is a block diagram of an electronic device provided by an embodiment of the present application. As Figure 13 shown, the electronic device includes, but is not limited to: a processor 1301 and a memory 1302.

[0335] Among them, the above-mentioned memory 1302 is used to store the executable instructions of the above-mentioned processor 1301. It can be understood that the above-mentioned processor 1301 is configured to execute instructions to implement the vehicle control method in the above-mentioned embodiment.

[0336] It should be noted that those skilled in the art can understand that Figure 13 the structure of the electronic device shown in Figure 13 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than

[0337] The processor 1301 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1302, and by invoking the data stored in the memory 1302, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 1301 may include one or more processing units. Optionally, the processor 1301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1301 either.

[0338] The memory 1302 can be used to store software programs and various data. The memory 1302 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required by at least one functional module (such as a determination unit, a processing unit, etc.), etc. In addition, the memory 1302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0339] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as the memory 1302 including instructions. The above instructions can be executed by the processor 1301 of the electronic device to implement the method in the above embodiment.

[0340] In actual implementation, Figure 12 the functions of the acquisition unit 1201, the first determination unit 1202, the second determination unit 1203, and the processing unit 1204 in Figure 13 can all be implemented by the processor 1301 in

[0341] invoking the computer program stored in the memory 1302. The specific execution process can refer to the description of the method part in the above embodiment, and will not be elaborated here.

[0342] In an exemplary embodiment, the embodiment of the present application also provides a computer program product including one or more instructions. The one or more instructions can be executed by the processor 1301 of the electronic device to complete the method in the above embodiment.

[0343] It should be noted that when one or more instructions in the above-mentioned computer-readable storage medium or in the computer program product are executed by a processor of an electronic device, the various processes of the above method embodiments are implemented, and the same technical effects as those of the above method can be achieved. To avoid repetition, details are not described herein again.

[0344] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0345] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0346] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0347] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0348] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0349] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a first data frame, where the first data frame includes multiple data frames collected by a target vehicle at a first moment; When parameters included in the multiple data frames meet a first judgment condition, determining adjacent data frames of the first data frame based on the first moment and a time domain radius threshold, and adding the first data frame and the adjacent data frames of the first data frame to a domain data frame cluster; Determining a vector corresponding to each data frame in the domain data frame cluster, and determining a second data frame based on the number of data frames included in the domain data frame cluster and the vectors corresponding to each data frame, where each data frame includes continuous parameters and / or discrete parameters; Determining a target data frame with the highest similarity to the second data frame from the first data frame; Using the target data frame as the data frame at the first moment.

2. The method according to claim 1, wherein The first judgment condition includes: The range of at least one continuous parameter included in the multiple data frames is greater than a range threshold; and / or, The discrete parameters included in the multiple data frames are different.

3. The method according to claim 1 or 2, characterized in that, The method further includes: When parameters included in the multiple data frames meet a second judgment condition, selecting any data frame from the first data frame as the data frame at the first moment.

4. The method according to claim 3, wherein Each data frame includes continuous parameters and / or discrete parameters, and the second judgment condition includes: The range of each continuous parameter included in the multiple data frames is less than or equal to the range threshold, and each discrete parameter included in the multiple data frames is the same.

5. The method according to claim 1, characterized in that, The method further includes: When the number of adjacent data frames of the first data frame is less than a threshold, generating a third data frame based on the adjacent data frames of the first data frame; The determining the second data frame based on the adjacent data frames of the first data frame and the first data frame includes: Determining the second data frame based on the adjacent data frames of the first data frame, the first data frame, and the third data frame.

6. The method according to claim 1 or 2, characterized in that, The determining the target data frame with the highest similarity to the second data frame from the first data frame includes: For each data frame in the first data frame, determining the degree of difference between the data frame and the second data frame; the degree of difference is negatively correlated with the similarity; Using the data frame in the first data frame with the smallest degree of difference from the second data frame as the target data frame.

7. The method according to claim 1, characterized in that, The method further includes: Obtaining multiple data frames collected by the target vehicle at different moments; For each data frame in the multiple data frames collected by the target vehicle at different moments, determining the vehicle operating state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame; the vehicle state refers to the state of the vehicle in different working modes; the vehicle operating state includes different states during the operation of the vehicle. For any two adjacent data frames among multiple data frames collected for the target vehicle at different times, determine data frames belonging to the same driving segment based on the vehicle running states corresponding to the two adjacent data frames and the time interval between the two adjacent data frames. The obtaining of the first data frame includes: Obtain the first data frame from among the multiple data frames included in the same driving segment.

8. The method according to claim 7, wherein The vehicle running states include a first running state and a second running state. The first running state means that the vehicle state is the starting state, and the charging state is the driving charging state or the non-charging state. The driving charging state means that during vehicle driving, the battery is in the energy recovery state. The second running state means that the vehicle state is in the non-starting state.

9. The method according to claim 7 or 8, characterized in that, Before determining the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame, the method further includes: For each data frame included in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meet a preset correction condition, correct the vehicle state and / or charging state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame.

10. The method according to claim 9, wherein The correction condition includes any one of the following: The vehicle state corresponding to each data frame included in the same driving segment is another vehicle state, and the other vehicle state means a vehicle state other than the starting state and the extinguished state. The vehicle state corresponding to each data frame included in the same driving segment is the starting state and the charging state is another charging state, and the other charging state means a charging state other than the driving charging state and the non-charging state.

11. The method according to claim 1, wherein The method further includes: For each data frame included in the first data frame, when the parameter included in the data frame is a missing value, successively use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained. The missing value includes a parameter with a null value or a parameter with a value exceeding a preset range. Among them, the order of the sub-imputation models for performing imputation processing on the missing value is related to the accuracy index of the sub-imputation model, and the accuracy index is used to characterize the accuracy rate of the imputation result obtained using the sub-imputation model.

12. A data processing device, characterized in that, The device includes: an obtaining unit, a first determining unit, a second determining unit, and a processing unit; The obtaining unit is used to obtain a first data frame, and the first data frame includes multiple data frames collected for the target vehicle at a first time. The first determining unit is used for: When the parameters included in the multiple data frames meet a first judgment condition, determine the adjacent data frames of the first data frame based on the first time and the time domain radius threshold, and add the first data frame and the adjacent data frames of the first data frame to the domain data frame cluster. Determine the vectors corresponding to each data frame in the domain data frame cluster, and determine the second data frame based on the number of data frames included in the domain data frame cluster and the vectors corresponding to each data frame. Each of the data frames includes continuous parameters and / or discrete parameters; The second determination unit is configured to determine, from the first data frames, a target data frame that has the highest similarity to the second data frame; The processing unit is configured to use the target data frame as the data frame at the first moment.

13. An electronic device, characterized in that, It includes a memory and a processor; The memory is coupled to the processor; The memory is used to store computer program code, and the computer program code includes computer instructions; When the processor executes the computer instructions, the electronic device executes the data processing method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, When the computer execution instructions stored in the computer-readable storage medium are executed by the processor of the processing device, the processing device can execute the data processing method according to any one of claims 1-11.

15. A computer program product, characterized in that, The computer program product includes the computer program, and the computer program is suitable for being loaded and executed by the processor to execute the data processing method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Image processing unit, image pickup device, image processing program, and image processing method

    CN106852190A

  • Video compression method with low loss

    CN111901600A