Data processing method and device, electronic equipment and storage medium

By determining the appropriate representative data frames in the car driving data, the problems of abnormal, duplication and missing in the data are solved, and high-quality processing and accuracy of the data are achieved.

CN120045860AActive Publication Date: 2025-05-27CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510511062.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

There are quality problems such as abnormality, duplication and missing in car driving data. The existing data preprocessing technology cannot effectively handle these problems, resulting in low data accuracy and reduced amount of available data.

Method used

By acquiring the target data frames in multiple data frames, determining the appropriate data frame as representative data frames based on the similarity and extreme difference threshold values, removing data frames with low similarity and improving data quality.

Benefits of technology

Ensure the accuracy and reliability of the ultimately retained target data frame, improve the overall quality of vehicle data, and reduce errors and waste in data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045860A_ABST
    Figure CN120045860A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method and device, electronic equipment and a storage medium, and relates to the field of data processing. The method comprises the steps that a first data frame is obtained, the first data frame comprises a plurality of data frames collected by a target vehicle at a first moment, then a second data frame is determined based on adjacent data frames of the first data frame and the first data frame, the adjacent data frames of the first data frame comprise data frames collected by the target vehicle at adjacent moments of the first moment, and the second data frame comprises a plurality of data frames collected by the target vehicle at adjacent moments of the first moment. And then, the target data frame with the highest similarity with the second data frame is determined from the first data frames, and the target data frame is taken as the data frame at the first moment, so that the quality problems of abnormity or repetition and the like in the driving data can be effectively solved, and the high-quality driving data can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Automobile driving data covers information such as driving behavior characteristics, energy consumption characteristics, and vehicle component working characteristics. This data not only provides key support for breakthroughs in core technologies such as vehicle performance optimization and battery health management, but also has strategic value for the construction of industrial ecosystems such as charging and swapping infrastructure planning and intelligent transportation system construction.

[0003] However, limited by hardware costs such as in-vehicle sensors, communication network devices, and storage devices, as well as multiple factors such as complex vehicle working conditions interference, the time-series driving data of automobiles generally has quality problems such as anomalies, repetitions, and omissions. These problems seriously interfere with the mining and research of the information contained in the driving data, and greatly hinder the transformation of data into practical results.

[0004] Currently, existing data preprocessing technologies do not comprehensively handle the possible quality defects in driving data. In addition to conventional outliers and missing values, the processing of duplicate data is often ignored. Moreover, when dealing with outliers and missing values, simple deletion methods or smoothing methods are often used, which inevitably cause problems such as a decrease in the amount of available data and low data accuracy. Based on this, it is of great significance to construct a driving data preprocessing method that can comprehensively handle the quality problems of driving data and ensure high data accuracy and small loss of available data volume. Summary of the Invention

[0005] One of the objectives of the present invention is to provide a data processing method, apparatus, electronic device, and storage medium, which can effectively solve the quality problems such as anomalies and repetitions in driving data and obtain high-quality driving data.

[0006] To achieve the above objective, the technical solution adopted by the present invention is as follows: According to the first aspect provided by the present invention, a data processing method is provided. The method includes: obtaining a first data frame, where the first data frame includes a plurality of data frames collected by a target vehicle at a first moment. Then, based on the adjacent data frames of the first data frame and the first data frame, a second data frame is determined. The adjacent data frames of the first data frame include the data frames collected by the target vehicle at adjacent moments of the first moment. Next, the target data frame with the highest similarity to the second data frame is determined from the first data frame, and the target data frame is used as the data frame at the first moment.

[0007] According to the above technical means, when there are data frames with the same acquisition time in the data of the target vehicle, the second data frame can be determined based on the data frames with the same acquisition time (i.e., the first data frame) and the adjacent data frames of the first data frame. The second data frame obtained in this way can comprehensively reflect the running information of the target vehicle at multiple moments. Then, the target data frame with the highest similarity to the second data frame in the first data frame is used as the data frame at the first moment, which can effectively remove the data frames with low similarity in the first data frame. In this way, not only can the accuracy and reliability of the finally retained target data frame be ensured, but also the overall quality of vehicle data can be improved.

[0008] In a possible way, to determine the second data frame based on the adjacent data frames of the first data frame and the first data frame, it may specifically include: when the parameters included in multiple data frames meet the first judgment condition, determining the second data frame based on the adjacent data frames of the first data frame and the first data frame. Each data frame includes continuous parameters and / or discrete parameters.

[0009] According to the above technical means, the present application can perform the subsequent operation of determining the second data frame only when the parameters in multiple data frames meet the first judgment condition, avoiding the ineffective processing of data that does not meet specific requirements, and greatly improving the accuracy of data processing.

[0010] In a possible way, the first judgment condition includes: the range of at least one continuous parameter included in multiple data frames is greater than the range threshold; and / or, the discrete parameters included in multiple data frames are different.

[0011] According to the above technical means, the present application can quickly determine the data frames that meet the first judgment condition based on the above first judgment condition.

[0012] In a possible way, the above method further includes: when the parameters included in multiple data frames meet the second judgment condition, selecting any data frame from the first data frame as the data frame at the first moment.

[0013] According to the above technical means, when the parameters of multiple data frames meet the second judgment condition, the present application can directly select any one from the first data frame as the data frame at the first moment, without performing complex operations such as similarity calculation and multi-data frame correlation analysis, reducing the demand for hardware computing resources.

[0014] In a possible way, the second judgment condition includes: the range of each continuous parameter included in multiple data frames is less than or equal to the range threshold, and each discrete parameter included in multiple data frames is the same.

[0015] According to the above technical means, the present application can quickly determine the data frame that meets the second judgment condition based on the above second judgment condition.

[0016] In a possible way, the above method further includes: when the number of adjacent data frames of the first data frame is less than a threshold, generating a third data frame based on the adjacent data frames of the first data frame.

[0017] On this basis, determining the second data frame based on the adjacent data frames of the first data frame and the first data frame may specifically include: determining the second data frame based on the adjacent data frames of the first data frame, the first data frame, and the third data frame.

[0018] According to the above technical means, when the number of adjacent data frames of the first data frame is insufficient, the present application can generate a third data frame in combination with the adjacent data frames of the first data frame. In the subsequent process of determining the second data frame, the third data frame is used as supplementary information and combined with the adjacent data frames of the first data frame and the first data frame, further optimizing the process of determining the second data frame, enabling the second data frame to integrate more dimensional information, and improving the overall efficiency of data utilization.

[0019] In a possible way, determining the target data frame with the highest similarity to the second data frame from the first data frame may specifically include: for each data frame in the first data frame, determining the degree of difference between the data frame and the second data frame, and the degree of difference is negatively correlated with the similarity. Then, the data frame with the smallest degree of difference between the first data frame and the second data frame is used as the target data frame.

[0020] In a possible way, the above method further includes: obtaining a plurality of data frames collected by the target vehicle at different times. Then, for each data frame in the plurality of data frames collected by the target vehicle at different times, determining the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame. Finally, for any two adjacent data frames in the plurality of data frames collected by the target vehicle at different times, determining the data frames belonging to the same driving segment based on the vehicle running states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames. Wherein, the vehicle state refers to the state of the vehicle in different working modes, and the vehicle running state includes different states during the running process of the vehicle.

[0021] On this basis, obtaining the first data frame may specifically include: obtaining the first data frame from the plurality of data frames included in the same driving segment.

[0022] According to the above technical means, the present application can determine the vehicle running state by comprehensively considering the vehicle state and charging state corresponding to each data frame. This method can more comprehensively and accurately reflect the actual running situation of the vehicle at different times. Then, based on the vehicle running state and time interval corresponding to any two adjacent data frames, it is determined whether they belong to the same driving segment. This method can reasonably divide the driving process of the vehicle, and then obtain data with clear travel boundaries. At the same time, accurate driving segment division is of great significance for analyzing aspects such as the driving behavior, energy consumption, and driving route of the vehicle. In addition, obtaining the first data frame from multiple data frames included in the same driving segment makes the data processing more targeted and can improve the data processing efficiency.

[0023] In a possible way, the vehicle running state includes a first running state and a second running state. The first running state means that the vehicle state is the starting state, and the charging state is the driving charging state or the non-charging state. The driving charging state means that during the vehicle driving process, the battery is in the energy recovery state. The second running state means the state when the vehicle state corresponding to the data frame is the non-starting state.

[0024] In a possible way, before determining the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame, the above method further includes: for each data frame included in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meet the preset correction conditions, based on the charging state, vehicle speed, and total current corresponding to each data frame, correct the vehicle state and / or charging state corresponding to each data frame.

[0025] According to the above technical means, the present application can perform correction based on multi-dimensional information such as the charging state, vehicle speed, and total current when the vehicle state and / or charging state meet the preset correction conditions, and can correct data gaps, anomalies, and inaccuracies caused by various factors (such as sensor errors, signal interference, etc.).

[0026] In a possible way, the correction conditions include any one of the following: the vehicle state corresponding to each data frame included in the same driving segment is other vehicle states, and the other vehicle states refer to vehicle states other than the starting state and the flameout state. The vehicle state corresponding to each data frame included in the same driving segment is the starting state and the charging state is other charging states, and the other charging states refer to charging states other than the driving charging state and the non-charging state.

[0027] According to the above technical means, the present application can specifically screen out the data frames that need to be corrected based on the above preset correction conditions, so as to improve the data processing efficiency and save computing resources and time costs.

[0028] In one possible way, the above method further includes: for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, successively using the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained. The missing value includes a parameter with a null value or a parameter with a value exceeding a preset range.

[0029] Among them, the order of the sub-imputation models for performing imputation processing on the missing value is related to the accuracy index of the sub-imputation model. The accuracy index is used to characterize the accuracy rate of the imputation result obtained by using the sub-imputation model.

[0030] According to the above technical means, the present application can determine the imputation order according to the accuracy index of the sub-imputation model, and preferentially use the sub-imputation model with higher accuracy for imputation, so as to ensure the accuracy rate of the imputation result to the greatest extent. In addition, compared with directly deleting the data frame containing the missing value, using the imputation model for processing can retain more valid data and reduce data waste.

[0031] According to the second aspect provided by the present invention, there is provided a data processing device, which includes: an acquisition unit, a first determination unit, a second determination unit, and a processing unit.

[0032] The acquisition unit is used to acquire a first data frame, and the first data frame includes a plurality of data frames collected by the target vehicle at the first moment.

[0033] The first determination unit is used to determine a second data frame based on the adjacent data frame of the first data frame and the first data frame. The adjacent data frame of the first data frame includes the data frames collected by the target vehicle at the adjacent moment of the first moment.

[0034] The second determination unit is used to determine the target data frame with the highest similarity to the second data frame from the first data frame.

[0035] The processing unit is used to use the target data frame as the data frame at the first moment.

[0036] In one possible way, the first determination unit is further used to determine the second data frame based on the adjacent data frame of the first data frame and the first data frame when the parameters included in the plurality of data frames meet the first judgment condition.

[0037] In one possible way, the first determination unit is further used to generate a third data frame based on the adjacent data frame of the first data frame when the number of adjacent data frames of the first data frame is less than the threshold. Then, based on the adjacent data frame of the first data frame, the first data frame, and the third data frame, determine the second data frame.

[0038] In a possible way, the second determination unit is configured to, for each data frame in the first data frames, determine the degree of difference between the data frame and the second data frame, and use the data frame with the smallest degree of difference between the first data frames and the second data frame as the target data frame.

[0039] In a possible way, the acquisition unit is further configured to acquire a plurality of data frames collected by the target vehicle at different times.

[0040] In a possible way, the first determination unit is further configured to, for each data frame in the plurality of data frames collected by the target vehicle at different times, determine the vehicle running state corresponding to the data frame based on the vehicle state and the charging state corresponding to the data frame.

[0041] In a possible way, the first determination unit is further configured to, for any two adjacent data frames in the plurality of data frames collected by the target vehicle at different times, determine the data frames belonging to the same driving segment based on the vehicle running states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames.

[0042] In a possible way, the acquisition unit is further configured to acquire the first data frame from the plurality of data frames included in the same driving segment.

[0043] In a possible way, the processing unit is further configured to, for each data frame included in the same driving segment, when the vehicle state and / or the charging state corresponding to each data frame meet the preset correction conditions, correct the vehicle state and / or the charging state corresponding to each data frame based on the charging state, the vehicle speed, and the total current corresponding to each data frame.

[0044] In a possible way, the processing unit is further configured to, for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, sequentially use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained.

[0045] According to a third aspect of the present invention, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement the method according to the first aspect and any possible implementation manner thereof.

[0046] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the processing device, enabling the processing device to execute the method according to the first aspect and any possible implementation manner thereof.

[0047] According to the fifth aspect provided by the present invention, there is provided a computer program product, which includes computer instructions. When the computer instructions run on a processing device, the processing device is caused to execute the method according to the first aspect and any possible implementation manner thereof described above.

[0048] Thus, the above technical features of the present invention have the following beneficial effects: (1) When there are data frames with the same acquisition time in the data of the target vehicle, according to the data frames with the same acquisition time (i.e., the first data frames) and the adjacent data frames of the first data frames, the second data frames are determined. The second data frames obtained in this way can comprehensively reflect the running information of the target vehicle at multiple times. Then, the target data frame with the highest similarity to the second data frame in the first data frame is used as the data frame at the first time, which can effectively remove the data frames with low similarity in the first data frame. In this way, not only can the accuracy and reliability of the finally retained target data frames be ensured, but also the overall quality of the vehicle data can be improved.

[0049] (2) The operation of determining the second data frame can be performed only when the parameters in multiple data frames meet the first judgment condition, avoiding the ineffective processing of data that does not meet specific requirements, and greatly improving the accuracy of data processing.

[0050] (3) Based on the above first judgment condition, the data frames that meet the first judgment condition can be quickly determined.

[0051] (4) When the parameters of multiple data frames meet the second judgment condition, any one of the first data frames can be directly selected as the data frame at the first time, without performing complex operations such as similarity calculation and multi-data frame correlation analysis, reducing the demand for hardware computing resources.

[0052] (5) Based on the above second judgment condition, the data frames that meet the second judgment condition can be quickly determined.

[0053] (6) When the number of adjacent data frames of the first data frame is insufficient, the third data frame can be generated by combining the adjacent data frames of the first data frame. In the subsequent process of determining the second data frame, the third data frame is used as supplementary information and combined with the adjacent data frames and the first data frame of the first data frame, further optimizing the process of determining the second data frame, enabling the second data frame to integrate more dimensional information, and improving the overall efficiency of data utilization.

[0054] (7) The vehicle operating state can be determined by comprehensively considering the vehicle state and charging state corresponding to each data frame. This method can more comprehensively and accurately reflect the actual operating conditions of the vehicle at different times. Subsequently, based on the vehicle operating state and time interval corresponding to any two adjacent data frames, it is determined whether they belong to the same driving segment. This method can reasonably divide the driving process of the vehicle, thereby obtaining data with clear trip boundaries. At the same time, accurate driving segment division is of great significance for analyzing aspects such as the driving behavior, energy consumption, and driving route of the vehicle. In addition, obtaining the first data frame from multiple data frames included in the same driving segment makes the data processing more targeted and can improve the data processing efficiency.

[0055] (8)When the vehicle state and / or charging state meet the preset correction conditions, corrections can be made based on multi-dimensional information such as the charging state, vehicle speed, and total current, which can correct data gaps, anomalies, and inaccuracies caused by various factors (such as sensor errors, signal interference, etc.).

[0056] (9)Based on the above preset correction conditions, data frames that need to be corrected can be selectively screened, which can improve the efficiency of data processing and save computing resources and time costs.

[0057] (10)The interpolation order can be determined according to the accuracy index of the sub-interpolation model, and the sub-interpolation model with higher accuracy is preferably used for interpolation, which can ensure the accuracy of the interpolation result to the greatest extent. In addition, compared with directly deleting data frames containing missing values, using the interpolation model for processing can retain more valid data and reduce data waste. Description of the Drawings

[0058] Figure 1 It is an architecture diagram of a data processing system provided by an embodiment of the present application; Figure 2 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 1 ; Figure 3 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 2 ; Figure 4 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 3 ; Figure 5 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 4 ; Figure 6 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 5 ; Figure 7 Flow schematic of a data processing method provided by an embodiment of the present application Figure 6 ; Figure 8 Flow schematic of a data processing method provided by an embodiment of the present application Figure 7 ; Figure 9 Flow schematic of a data processing method provided by an embodiment of the present application Figure 8 ; Figure 10 Flow schematic diagram of a method for constructing a target integrated interpolation model provided by an embodiment of the present application; Figure 11 Flow schematic of a data processing method provided by an embodiment of the present application Figure 9 ; Figure 12 Structural schematic diagram of a data processing device provided by an embodiment of the present application; Figure 13 Block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0059] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0060] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0061] In the embodiments of the present application, words such as "exemplary", "for example", or "such as" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "such as" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "such as" is intended to present related concepts in a specific manner.

[0062] First, the related technologies involved in the present application are explained and described to facilitate the understanding of those skilled in the art.

[0063] Driven by the global energy structure transformation and the need for low-carbon development, electric vehicles have become the strategic direction of the automotive industry innovation with their core advantages of high energy efficiency and low emissions. With the continuous breakthroughs in battery, motor, and electronic control technologies and the in-depth application of intelligent network technology, the real-time operation data of electric vehicles has shown an explosive growth trend. The vehicle supervision big data platform built in accordance with relevant technical specifications has effectively opened up the collection, transmission and storage links of vehicle driving data, forming a massive data set covering driving behavior characteristics, energy consumption characteristics, and vehicle component working characteristics. These data not only provide key support for core technology breakthroughs such as vehicle performance optimization and battery health management, but also have strategic value for the construction of industrial ecology such as charging and swapping infrastructure planning and intelligent transportation system construction.

[0064] However, due to the cost constraints of on-board sensors, communication network equipment, storage devices and other hardware, as well as the influence of multiple factors such as interference from complex vehicle working conditions, the original electric vehicle time series driving data generally has quality problems such as anomalies, duplications, and missing data, and there is a lack of clear boundaries for each trip data. These problems seriously interfere with the mining of information contained in driving data and greatly hinder the transformation of data into practical results.

[0065] At present, there are many deficiencies in the relevant technologies related to vehicle driving data preprocessing. On the one hand, the existing preprocessing solutions are mostly developed for the private protocol data of specific automobile companies, which are difficult to apply to driving data collected in accordance with the common standard GB / T32960 and lack universality; on the other hand, the processing of possible quality defects in driving data is not comprehensive enough, and in addition to conventional outliers and missing values, the processing of duplicate data is often ignored. In addition, when processing outliers and missing values, simple deletion methods or time series interpolation and smoothing methods are often used, which inevitably leads to problems such as a decrease in the amount of available data and low data accuracy.

[0066] Based on this, there is an urgent need to build a driving data preprocessing method that is highly versatile, can comprehensively deal with driving data quality issues, and ensure high data accuracy and small loss of available data volume, so as to build a high-quality data base for automobile-related research and ensure the authenticity, effectiveness and stability of its results.

[0067] To solve the above technical problems, an embodiment of the present application provides a data processing method. When there are data frames with the same acquisition time in the data of the target vehicle, the second data frame can be determined according to the data frames with the same acquisition time (i.e., the first data frame) and the adjacent data frames of the first data frame. The second data frame obtained in this way can comprehensively reflect the running information of the target vehicle at multiple moments. Then, the target data frame with the highest similarity to the second data frame in the first data frame is used as the data frame at the first moment, which can effectively remove the data frames with low similarity in the first data frame. In this way, not only can the accuracy and reliability of the finally retained target data frames be ensured, but also the overall quality of the vehicle data can be improved.

[0068] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0069] Figure 1 It is an architecture diagram of a data processing system provided by an embodiment of the present application, as Figure 1 shown. The system architecture includes: a server 101.

[0070] Among them, the server 101 can be a high-performance server that provides various services on the Internet. It can be an independent physical server, or a server cluster composed of multiple physical servers, or at least one of the cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks, and big data or artificial intelligence platforms. The embodiments of the present application do not limit this. Of course, the server can also include other functions to provide more comprehensive and diverse services.

[0071] The server 101 in the embodiments of the present application can be a single server, a server cluster, or a cloud server. The embodiments of the present application do not limit this.

[0072] In the embodiments of the present application, when there are multiple data frames collected at the same moment in the data of the target vehicle, the server 101 can obtain the first data frame, where the first data frame includes multiple data frames collected by the target vehicle at the first moment. Then, the server 101 can determine the second data frame based on the adjacent data frames and the first data frame of the first data frame, and determine the similarity between each data frame in the first data frame and the second data frame. Then, the server 101 can use the data frame with the highest similarity to the second data frame in the first data frame as the data frame at the first moment.

[0073] The data processing system provided by the embodiments of this application can be configured in a vehicle. A vehicle can also be referred to as a transportation vehicle (vehicle), a mobile carrier, an electric vehicle (EV), a hybrid electric vehicle (HEV), a plug-in hybrid electric vehicle (PHEV), a fuel cell vehicle (FCV), an autonomous vehicle, an intelligent and connected vehicle (ICV), a driverless vehicle, etc.

[0074] In the embodiments of this application, the vehicle can be a sedan, a sport utility vehicle (SUV), a truck, an electric vehicle, a motorcycle, a tricycle, a special vehicle (such as an ambulance, a fire truck, a police car, etc.), a driverless taxi, an intelligent and connected bus, an autonomous logistics vehicle, an electric truck, etc. In addition, this method is also applicable to various special vehicles, such as agricultural vehicles, mining vehicles, forestry vehicles, airport vehicles, port vehicles, etc. This application does not make specific limitations in this regard.

[0075] For ease of understanding, the data processing method provided by this application will be specifically introduced below with reference to the accompanying drawings.

[0076] Figure 2 It is a schematic flowchart of a data processing method provided by the embodiments of this application. As Figure 2 shown, this method is executed by the Figure 1 server shown, and this method includes: S201. Obtain a first data frame.

[0077] Among them, the first data frame includes multiple data frames collected by the target vehicle at the first moment, that is, the collection times of the multiple data frames are the same.

[0078] In some embodiments, the server stores data frames collected by the target vehicle at different moments. On this basis, the server can screen out multiple data frames collected at the first moment from the data frames collected at different moments, so as to obtain the first data frame including multiple data frames with the same collection time.

[0079] S202. Determine a second data frame based on the adjacent data frames of the first data frame and the first data frame.

[0080] Among them, the adjacent data frames of the first data frame include the data frames collected by the target vehicle at adjacent moments of the first moment. In the embodiments of the present application, the adjacent moments of the first moment may be moments adjacent or close to the first moment. For example, if the first moment is 10 o'clock, the adjacent moments of the first moment may be 10:01 or 10:05, and there is no limitation thereto.

[0081] In some embodiments, the server may determine the adjacent data frames of the first data frame based on the first moment and the time domain radius threshold. After that, the server may add the first data frame and the adjacent data frames of the first data frame to the domain data frame cluster (i.e., the data frame set), and determine the vector corresponding to each data frame in the domain data frame cluster. Then, the server may determine the second data frame based on the number of data frames included in the domain data frame cluster and the vectors corresponding to each data frame.

[0082] In the embodiments of the present application, the vector corresponding to the data frame is the vector composed of the parameters included in the data frame.

[0083] Exemplarily, taking the first moment as t and the time domain radius threshold as as an example. The server adds the data frames with the acquisition moment in the range to the domain data frame cluster, and the other data frames in the domain data frame cluster except the first data frame are the adjacent data frames of the first data frame. After that, the server may determine the second data frame with reference to the following formula 1.

[0084] (Formula 1).

[0085] Among them, is the second data frame (which may also be referred to as the center of the domain data frame cluster), is the number of data frames included in the domain data frame cluster, is the vector corresponding to the i-th data frame in the domain data frame cluster.

[0086] S203. Determine the target data frame with the highest similarity to the second data frame from the first data frame.

[0087] In some embodiments, for each data frame in the first data frame, the server may determine the degree of difference between the data frame and the second data frame. After that, the server may use the data frame with the smallest degree of difference between the first data frame and the second data frame as the target data frame.

[0088] Among them, the degree of difference is negatively correlated with the similarity. For example, the smaller the degree of difference between any data frame in the first data frame and the second data frame, the higher the similarity between the data frame and the second data frame, and the larger the degree of difference between any data frame in the first data frame and the second data frame, the lower the similarity between the data frame and the second data frame.

[0089] The embodiments of the present application do not limit the way to represent the degree of difference. For example, the Minkowski distance (such as the Euclidean distance, Manhattan distance, etc.) can be used to represent the degree of difference, or the Mahalanobis distance can be used to represent the degree of difference, which is not limited herein. Among them, the distance is positively correlated with the degree of difference, that is, the smaller the distance, the smaller the degree of difference, and the larger the distance, the larger the degree of difference.

[0090] Exemplarily, taking the use of the Mahalanobis distance to represent the degree of difference as an example. The server can calculate the Mahalanobis distance between each data frame in the first data frame and the second data frame, and then use the data frame in the first data frame with the smallest Mahalanobis distance from the second data frame as the target data frame. Among them, the calculation formula of the Mahalanobis distance can refer to the following formula 2.

[0091] (Formula 2).

[0092] Among them, is the Mahalanobis distance from the data frame to the second data frame, is the vector corresponding to the data frame, is the covariance matrix corresponding to the domain data frame cluster, The meaning of can refer to the above formula 1 and will not be elaborated herein.

[0093] S204. Use the target data frame as the data frame at the first moment.

[0094] In some embodiments, the server can retain the target data frame, that is, use the target data frame as the data frame of the target vehicle at the first moment, and delete other data frames in the first data frame except the target data frame.

[0095] Based on the above technical means, when there are data frames with the same acquisition time in the data of the target vehicle, the second data frame can be determined according to multiple data frames with the same acquisition time (i.e., the first data frame) and the adjacent data frames of the first data frame. In this way, the obtained second data frame can comprehensively reflect the running information of the target vehicle at multiple moments. Then, use the target data frame with the highest similarity to the second data frame in the first data frame as the data frame at the first moment, and delete other data frames in the first data frame except the target data frame. In this way, not only can the accuracy and reliability of the finally retained target data frame be ensured, but also the data duplication problem can be solved, and the overall quality of vehicle data can be improved.

[0096] In the embodiments of the present application, each data frame included in the first data frame may include different parameters. The following takes (1) the parameters included in multiple data frames satisfy the first judgment condition and (2) the parameters included in multiple data frames satisfy the second judgment condition as examples to introduce the data processing method provided by the present application.

[0097] (1) The parameters included in multiple data frames satisfy the first judgment condition.

[0098] In an alternative embodiment, the above S202 may include: when the parameters included in multiple data frames satisfy the first judgment condition, the server may determine a second data frame based on the adjacent data frames and the first data frame of the first data frame.

[0099] Among them, each data frame may include continuous parameters and / or discrete parameters. A continuous parameter refers to a parameter that can take any value within a target interval, and a discrete parameter refers to a parameter whose value is finite or countable.

[0100] In the embodiments of the present application, the discrete parameters may include, but are not limited to, the positioning state of the vehicle, the vehicle state, and the charging state. The continuous parameters may include, but are not limited to, the drive motor speed, the drive motor torque, the drive motor temperature, the input voltage of the motor controller, the DC bus current of the motor controller, the longitude, the latitude, the vehicle speed, the cumulative mileage, the total voltage, the total current, and the state of charge (SOC).

[0101] On this basis, the first judgment condition includes at least one of the following: 1-1. The range of at least one continuous parameter included in multiple data frames is greater than the range threshold.

[0102] In the embodiments of the present application, the range threshold corresponding to the continuous parameter may be determined according to actual needs, and no limitation is made thereto. For example, Table 1 shows the range thresholds of the continuous parameters included in the data frames.

[0103] Table 1 Range Thresholds of Continuous Parameters Included in Data Frames

[0104] Exemplarily, in combination with Table 1 above, taking speed as an example. Assume that the speed corresponding to Data Frame 1 is 3 km / h, the speed corresponding to Data Frame 2 is 6 km / h, and the speed corresponding to the data frame is 4 km / h. The range of these three data frames is the difference (3 km / h) between the maximum value (6 km / h) and the minimum value (3 km / h). Since the range of the three data frames (3 km / h) is greater than the speed range threshold (2 km / h), it indicates that the first judgment condition is satisfied.

[0105] 1-2. The discrete parameters included in multiple data frames are different.

[0106] Exemplarily, taking the vehicle state as an example, the vehicle state may include a starting state, a shutdown state, and other states. On this basis, if the vehicle state corresponding to data frame 1 is the starting state, the vehicle state corresponding to data frame 2 is the shutdown state, and the vehicle state corresponding to data frame 3 is the starting state, it indicates that the first judgment condition is met.

[0107] In some embodiments, when the parameters included in multiple data frames meet any of the above first judgment conditions, the server may determine the second data frame with reference to the method in S202 above, which will not be elaborated here.

[0108] (2) The parameters included in multiple data frames meet the second judgment condition.

[0109] In an alternative embodiment, when the parameters included in multiple data frames meet the second judgment condition, the server may select any data frame from the first data frames as the data frame at the first moment.

[0110] Among them, the second judgment condition includes that the range of each continuous parameter included in multiple data frames is less than or equal to the range threshold, and each discrete parameter included in multiple data frames is the same.

[0111] In some embodiments, when the parameters included in multiple data frames meet the second judgment condition, the server may randomly retain any data frame and delete other data frames in the first data frames except this data frame.

[0112] Exemplarily, taking the continuous parameter as the vehicle speed and the discrete parameters as the vehicle state and the charging state as an example. If the range of the vehicle speeds corresponding to data frame 1, data frame 2, and data frame 3 is less than the vehicle speed range threshold, and the vehicle states and charging states corresponding to data frame 1, data frame 2, and data frame 3 are exactly the same, it indicates that the second judgment condition is met. In this case, the server may retain data frame 1 and delete data frame 2 and data frame 3.

[0113] Based on the above technical means, the present application can determine different duplicate data processing strategies based on the preset conditions met by the parameters included in multiple data frames. When the parameters in multiple data frames meet the first judgment condition, the operation of determining the second data frame is performed, avoiding ineffective processing of data that does not meet specific requirements, and greatly improving the accuracy of data processing. When the parameters of multiple data frames meet the second judgment condition, any one is directly selected from the first data frames as the data frame at the first moment, without other complex processing, reducing the demand for hardware computing resources.

[0114] In an alternative embodiment, before performing the above S202, the data processing method provided by the present application may further include: when the number of adjacent data frames of the first data frame is less than a threshold, the server may generate a third data frame based on the adjacent data frames of the first data frame. On this basis, the above S202 may be implemented as: the server may determine the second data frame based on the adjacent data frames of the first data frame, the first data frame, and the third data frame.

[0115] In some embodiments, when the number of adjacent data frames of the first data frame is less than a threshold, the server may use a preset oversampling method to generate a third data frame based on the adjacent data frames of the first data frame. The third data frame includes a plurality of new data frames generated by the preset oversampling method. After that, the server may refer to the method in the above S202 to determine the second data frame based on the first data frame, the adjacent data frames of the first data frame, the third data frame, and the number of the first data frame, the adjacent data frames of the first data frame, and the third data frame, which will not be elaborated here.

[0116] In the embodiments of the present application, the preset oversampling method may include random linear interpolation with noise addition oversampling, random oversampling, generative adversarial network oversampling, etc., which are not limited herein.

[0117] In other embodiments, the server may add the first data frame and the adjacent data frames of the first data frame to the domain data frame cluster and calculate the covariance matrix corresponding to the domain data frame cluster. After that, the server may calculate the determinant of the covariance matrix, and when the absolute value of the determinant is less than the determinant threshold, use a preset oversampling method to generate a third data frame based on the adjacent data frames of the first data frame.

[0118] Exemplarily, taking random linear interpolation with noise addition oversampling as an example. The server may successively use random linear interpolation with noise addition oversampling to generate a new data frame and add it to the domain data frame cluster until the absolute value of the determinant of the covariance matrix corresponding to the domain data frame cluster is greater than the determinant threshold. The new data frame may be determined by the following formula 3.

[0119] (Formula 3).

[0120] Where is the new data frame generated by oversampling, is the random linear interpolation weight, ; is a randomly selected data frame from the adjacent data frames of the first data frame and is not the last data frame, is the data frame is the adjacent data frame of; is a random vector, and Each component in is uniformly distributed and randomly generated on; is the maximum absolute value vector of noise. It should be noted that the above random variables, such as , etc. are randomly regenerated in each round of oversampling.

[0121] The following will be combined with Figure 3 , taking random linear interpolation with noise addition and oversampling as an example, to introduce the method in the above embodiments. As Figure 3 shown, the method includes: S301 - S308.

[0122] S301. Determine that the acquisition time of the first data frame is t, and the time domain radius threshold is .

[0123] S302. Add the data frames whose acquisition times are within the range of to the domain data frame cluster, and determine the data frames in the domain data frame cluster except the first data frame as the centered - removed domain data frame cluster.

[0124] S303. Calculate the absolute value DA of the determinant of the covariance matrix corresponding to the domain data frame cluster.

[0125] S304. Judge whether DA is less than . If so, execute S305; if not, execute S306.

[0126] S305. Perform random linear interpolation with noise addition and oversampling within the centered - removed domain data frame cluster, generate new data frames and add them to the domain data frame cluster, and then execute S303.

[0127] S306. Determine the second data frame based on each data frame in the domain data frame cluster.

[0128] S307. Calculate the Mahalanobis distance between each data frame in the first data frame and the second data frame.

[0129] S308. Retain the data frame in the first data frame with the smallest Mahalanobis distance from the second data frame, and delete the other data frames in the first data frame.

[0130] Based on the above technical solution, when the number of adjacent data frames of the first data frame is insufficient, the present application can combine the adjacent data frames of the first data frame to generate a third data frame. Or, when the absolute value of the determinant of the covariance matrix corresponding to the first data frame and the adjacent data frames of the first data frame is less than the absolute value threshold, a third data frame is generated to avoid the determinant of the covariance matrix being too small to calculate the Mahalanobis distance between each data frame in the first data frame and the second data frame. In addition, during the subsequent determination of the second data frame, the third data frame is used as supplementary information and combined with the adjacent data frames of the first data frame and the first data frame to further optimize the determination process of the second data frame, enabling the second data frame to integrate more dimensional information and improving the overall efficiency of data utilization.

[0131] In an alternative embodiment, before performing S201 above, the data processing method provided by the present application may further include: The server can obtain multiple data frames collected by the target vehicle at different times. Then, for each data frame in the multiple data frames collected by the target vehicle at different times, the server can determine the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame. Finally, for any two adjacent data frames in the multiple data frames collected by the target vehicle at different times, the server can determine the data frames belonging to the same driving segment based on the vehicle running states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames. On this basis, S201 above may specifically include: The server can obtain the first data frame from the multiple data frames included in the same driving segment.

[0132] In the embodiments of the present application, the vehicle state refers to the state of the vehicle in different working modes, and the vehicle running state includes different states during the running process of the vehicle.

[0133] In some embodiments, the vehicle running state may include a first running state and a second running state. The vehicle state may include a starting state and a non-starting state. The charging state may include a driving charging state and a non-charging state. The driving charging state refers to the state when the battery recovers energy through the energy recovery system during the vehicle's driving process.

[0134] On the above basis, after the server obtains multiple data frames collected by the target vehicle at different times, for each data frame in the multiple data frames collected by the target vehicle at different times, the server can determine the vehicle running state corresponding to the data frame with the vehicle state being the starting state and the charging state being the driving charging state or the non-charging state as the first running state, and determine the vehicle running state corresponding to the data frame with the vehicle state being the non-starting state as the second running state.

[0135] After that, for any two adjacent data frames among multiple data frames collected for the target vehicle at different times, if the vehicle running states corresponding to any two adjacent data frames are different, then the two adjacent data frames do not belong to the same driving segment; if the vehicle running states corresponding to any two adjacent data frames are the same but the time interval between the two adjacent data frames is greater than the time interval threshold, then the two adjacent data frames do not belong to the same driving segment. After that, for the data frames in each driving segment, the server can delete the data frames in the second running state. Finally, the server can obtain the first data frame from the multiple data frames included in the same driving segment with reference to the method in S201 above, which will not be elaborated here.

[0136] In the embodiments of the present application, the time interval threshold can be 500 seconds, 600 seconds, etc., which is not limited herein.

[0137] In one example, the data frames in the first running state are defined as driving frames, and the data frames in the second running state are defined as non-driving frames. On this basis, the method for determining the data frames belonging to the same driving segment will be introduced in detail below in combination with the above embodiments, as Figure 4 shown, the method includes: S401 - S412.

[0138] S401. Sort the multiple data frames of the target vehicle according to the acquisition time.

[0139] Specifically, the server can arrange the multiple data frames of the target vehicle in ascending order of the acquisition time.

[0140] S402. Determine whether the initial data frame (i.e., the 0th data frame) of the target vehicle is a driving frame. If so, execute S403; if not, execute S404.

[0141] S403. Determine the initial data frame as the starting frame of the new driving segment.

[0142] S404. Initialize the frame index i = 1.

[0143] S405. Determine whether the frame index i is less than or equal to i_max. If so, execute S406; if not, execute S411.

[0144] Wherein, i_max is the number of multiple data frames.

[0145] S406. Determine whether the i-th data frame is a driving frame. If so, execute S407; if not, execute S410.

[0146] S407. Determine whether the (i - 1)-th data frame is a driving frame. If so, execute S408; if not, execute S409.

[0147] S408. Determine whether the time interval between the i-th data frame and the (i - 1)-th data frame is greater than the time interval threshold. If so, execute S409; if not, execute S410.

[0148] S409. Determine the i-th data frame as the starting frame of a new driving segment.

[0149] S410. Let the frame index i = i + 1.

[0150] S411. Determine the i_max-th data frame as the starting frame of a new segment.

[0151] S412. Determine the data frames between two adjacent starting frames of new segments, and the data frame with a lower order among two adjacent starting frames of new segments as the same driving segment.

[0152] Exemplarily, if data frame 1 is the starting frame of a new segment and data frame 4 is the starting frame of a new segment, then data frames 1, 2, and 3 are the same driving segment. After that, the server can also determine the driving segment numbers of each data frame. For example, the driving segment numbers of data frames 1, 2, and 3 are 1, and the driving segment number of data frame 4 is 2.

[0153] Based on the above technical means, the present application can determine the vehicle running state by comprehensively considering the vehicle state and charging state corresponding to each data frame. This method can more comprehensively and accurately reflect the actual running situation of the vehicle at different times. Then, based on the vehicle running state and time interval corresponding to any two adjacent data frames, determine whether they belong to the same driving segment. This method can reasonably divide the driving process of the vehicle, and then obtain data with clear travel boundaries. At the same time, accurate driving segment division is of great significance for analyzing the driving behavior, energy consumption, driving route, etc. of the vehicle. At the same time, by deleting the data frames in the second running state, the storage space can be saved and the efficiency of subsequent data processing can be improved. In addition, obtaining the first data frame from multiple data frames included in the same driving segment makes the data processing more targeted and can improve the data processing efficiency.

[0154] In an alternative embodiment, before determining the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame, the method provided by the present application may further include: for each data frame included in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meet the preset correction conditions, correct the vehicle state and / or charging state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame.

[0155] Among them, the correction conditions include any one of the following: 2-1. The vehicle state corresponding to each data frame included in the same driving segment is other vehicle states.

[0156] Among them, other vehicle states refer to vehicle states other than the starting state and the shutdown state.

[0157] 2-2. The vehicle state corresponding to each data frame included in the same driving segment is the starting state and the charging state is other charging states. Among them, other charging states refer to charging states other than the driving charging state and the non-charging state.

[0158] It should be noted that due to reasons such as equipment failures in the vehicle, there may be missing values, outliers, or meaningless values in the vehicle state and charging state corresponding to each data frame. Therefore, it is necessary to correct the vehicle state and charging state corresponding to each data frame.

[0159] Combined with the above content, the following takes (1) correcting the vehicle state corresponding to the data frame and (2) correcting the charging state corresponding to the data frame as examples for introduction.

[0160] (1) Correct the vehicle state corresponding to the data frame.

[0161] In some embodiments, for each data frame included in the same driving segment, when the vehicle state corresponding to each data frame is other vehicle states, the server can correct the vehicle state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame. Among them, the following explanations are made for the correction of the vehicle state: 1.1. For safety considerations, parking charging is usually carried out when the vehicle is shut down. Therefore, it is default that when the charging state is parking charging, the vehicle state is the shutdown state.

[0162] 1.2. When the charging state is the driving charging state, it indicates that the vehicle is in regenerative braking, and the vehicle state should be the starting state.

[0163] 1.3. When the total current is positive, it indicates that the power battery is outputting electrical energy, and the vehicle state should be the starting state.

[0164] Combined with the above explanations, taking the correction of the vehicle state corresponding to each data frame included in the target driving segment as an example, as Figure 5 shown, the method includes: S501-S513.

[0165] S501. Determine whether there is a data frame with a charging state of parking charging state in the target driving segment. If so, execute S502; if not, execute S503.

[0166] S502. Determine the vehicle state corresponding to each data frame included in the target driving segment as the shutdown state.

[0167] S503. Determine whether there is a data frame with a charging state of in - motion charging state in the target driving segment. If so, execute S504; if not, execute S505.

[0168] S504. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0169] S505. Determine whether more than half of the data frames in the target driving segment have a vehicle speed greater than 1 km / h. If so, execute S506; if not, execute S507.

[0170] S506. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0171] S507. Determine whether more than half of the data frames in the target driving segment have a vehicle speed less than or equal to 1 km / h. If so, execute S508; if not, execute S511.

[0172] S508. Determine whether more than half of the data frames in the target driving segment have a total current greater than 1 A. If so, execute S509; if not, execute S510.

[0173] S509. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0174] S510. Determine the vehicle state corresponding to each data frame included in the target driving segment as the off state.

[0175] S511. Determine whether more than half of the data frames in the target driving segment have a total current greater than 1 A. If so, execute S512; if not, execute S513.

[0176] S512. Determine the vehicle state corresponding to each data frame included in the target driving segment as the start state.

[0177] S513. Determine the vehicle state corresponding to each data frame included in the target driving segment as the invalid state.

[0178] In the embodiments of the present application, the invalid state can be represented by 255, or can be represented by other means, which is not limited herein.

[0179] (2) Correct the charging state corresponding to the data frame.

[0180] In some embodiments, for each data frame included in the same driving segment, when the vehicle state corresponding to each data frame is the starting state and the charging state is other charging states, the server may correct the charging state corresponding to each data frame based on the charging state, vehicle speed, and total current corresponding to each data frame. Among them, the following explanations are made for the correction of the charging state: 2.1. When the vehicle is in the starting state (for example, the vehicle speed corresponding to more than half of the data frames in the target driving segment is greater than 1 km / h), uniformly correct the charging state to the non-charging state.

[0181] 2.2. When the information obtained from the charging state, vehicle speed, and total current is contradictory or all are missing values, correct the charging state to 255, that is, the invalid state.

[0182] 2.3. If it is determined according to the vehicle speed, total current, and charging state that the vehicle is charging while parked in the starting state, respect the data record and allow the contradiction between the vehicle state and the charging state, and correct the charging state to the parked charging state.

[0183] 2.4. The charging completed state may occur both after the parked charging is full and at the initial stage of vehicle startup, and cannot be used to judge the driving state. When this value exists and the charging state cannot be determined through other parameters, the charging state remains the charging completed state.

[0184] Combined with the above explanations, taking the correction of the charging state corresponding to each data frame included in the target driving segment as an example, as Figure 6 shown, the method includes: S601 - S613.

[0185] S601. Judge whether there are more than half of the data frames in the target driving segment whose corresponding vehicle speeds are greater than 1 km / h. If so, execute S602; if not, execute S603.

[0186] S602. Determine the charging state corresponding to each data frame included in the target driving segment as the non-charging state.

[0187] S603. Judge whether there are more than half of the data frames in the target driving segment whose corresponding total currents are less than or equal to -1 A. If so, execute S604; if not, execute S607.

[0188] S604. Judge whether there is a data frame with a charging state of parked charging state in the target driving segment. If so, execute S605; if not, execute S606.

[0189] S605. Determine the charging state corresponding to each data frame included in the target driving segment as the parked charging state.

[0190] S606. Determine the charging status corresponding to each data frame included in the target driving segment as an invalid charging status.

[0191] S607. Determine whether there are more than half of the data frames in the target driving segment whose total current is greater than -1A. If so, execute S608; if not, execute S611.

[0192] S608. Determine whether there is a data frame with a charging status of fully charged in the target driving segment. If so, execute S609; if not, execute S610.

[0193] S609. Determine the charging status corresponding to each data frame included in the target driving segment as a fully charged status.

[0194] S610. Determine the charging status corresponding to each data frame included in the target driving segment as an uncharged status.

[0195] S611. Determine whether there is a data frame with a charging status of charging while parked in the target driving segment. If so, execute S612; if not, execute S613.

[0196] S612. Determine the charging status corresponding to each data frame included in the target driving segment as a charging while parked status.

[0197] S613. Determine the charging status corresponding to each data frame included in the target driving segment as an invalid charging status.

[0198] Based on the above technical means, the present application can correct problems such as value gaps, anomalies, and inaccuracies in the charging status and vehicle status caused by various factors such as sensor errors and signal interference, improve data quality, and provide support for subsequent determination of the vehicle operating status corresponding to each data frame.

[0199] In an optional implementation manner, the data processing method provided by the present application may further include: for each data frame included in the first data frame, when the parameter included in the data frame is a missing value, the server can successively use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained.

[0200] Among them, the missing value includes a parameter with a null value or a parameter with a value exceeding a preset range.

[0201] In the embodiments of the present application, each parameter included in the data frame corresponds to a threshold range (i.e., a reasonable range), and the threshold range is determined based on the specification parameters and common sense of the components in the vehicle to which the data frame belongs. For example, Table 2 shows the threshold ranges of each parameter. Referring to Table 2, for each parameter, the server can replace the data with a parameter value exceeding the corresponding threshold range with a missing value.

[0202] Table 2 Threshold ranges of each parameter

[0203] Among them, the order of the sub-imputation models for imputing missing values is related to the accuracy index of the sub-imputation models. The accuracy index is used to characterize the accuracy rate of the imputation results obtained by using the sub-imputation models.

[0204] The embodiments of the present application do not limit the sub-imputation models. For example, the sub-imputation models may include a mean imputation model, a random hot deck method, a regression imputation model, a time series imputation model, etc., and no limitation is made thereto. Among them, the regression imputation model may include a LightGBM model, a K-nearest neighbor model, a random forest, etc. The time series imputation model may include an adjacent frame linear interpolation model, a cubic Lagrange interpolation model, etc.

[0205] In the embodiments of the present application, the accuracy index may include a root mean squared error (RMSE), a mean squared error (MSE), etc., and no limitation is made thereto.

[0206] In some embodiments, multiple integrated imputation models are deployed in the server. Each integrated imputation model includes multiple sub-imputation models. Each parameter included in the data frame corresponds to an integrated imputation model. On this basis, taking the target parameter included in the target data frame in the first data frame as a missing value as an example, a method for imputing using the target integrated imputation model corresponding to the target parameter is introduced. As Figure 7 shown, the method includes: S701 - S707.

[0207] S701. Sort each sub-imputation model according to the accuracy index of each sub-imputation model in the target integrated imputation model.

[0208] Among them, the larger the accuracy index of the sub-imputation model, the higher the order of the sub-imputation model, that is, the sub-imputation model with a high accuracy index can be preferentially used for imputation processing.

[0209] S702. Initialize the sub-imputation model index i = 1.

[0210] S703. Determine whether the sub-imputation model index i is less than or equal to the total number i_max of sub-imputation models in the target integrated imputation model. If so, execute S704; if not, execute S707.

[0211] S704. Determine whether the other parameters in the target data frame except the target parameter meet the input parameters required by the i-th sub-imputation model. If so, execute S705; if not, execute S706.

[0212] S705. Interpolate the target parameter using the i-th sub-interpolation model.

[0213] S706. Let the sub-interpolation model index i = i + 1, and execute S703.

[0214] S707. Use other interpolation models other than the sub-interpolation models in the integrated interpolation model to interpolate the target parameter.

[0215] In the embodiments of the present application, other interpolation models may include a wide-range linear interpolation model, a nearest neighbor interpolation model, etc., which are not limited herein.

[0216] In one example, taking the longitude and latitude in the target data frame as missing values, the server can use the wide-range linear interpolation model to interpolate the longitude and latitude to obtain an interpolation result. If the interpolation using the wide-range linear interpolation model fails, the server can use the nearest neighbor interpolation model to interpolate the longitude and latitude.

[0217] In another example, taking the cumulative mileage in the target data frame as a missing value, the server can determine the interpolation value of the cumulative mileage corresponding to the target data frame based on the cumulative mileage corresponding to the previous data frame adjacent to the target data frame, the target data frame, the vehicle speed corresponding to the previous data frame adjacent to the target data frame, and the acquisition time corresponding to the previous data frame adjacent to the target data frame. For example, the interpolation value of the cumulative mileage corresponding to the target data frame satisfies the following formula 4.

[0218] (Formula 4).

[0219] Wherein, is the interpolation value of the cumulative mileage corresponding to the target data frame, is the cumulative mileage corresponding to the previous data frame adjacent to the target data frame, is the acquisition time of the target data frame, is the acquisition time of the previous data frame adjacent to the target data frame, is the vehicle speed corresponding to the target data frame, is the vehicle speed corresponding to the previous data frame adjacent to the target data frame. If is greater than a predetermined threshold (such as 1 / 240 h), or is a missing value, the nearest neighbor interpolation value can be used to interpolate the cumulative mileage corresponding to the target data frame.

[0220] Based on the above technical means, the sub-imputation model with high precision indicators can be preferentially used to impute missing values. When there are missing values in the input parameters and the current sub-imputation model cannot be used, the next sub-imputation model in the integrated imputation model is used sequentially until the imputation result is obtained. In this way, high-precision missing value imputation can be achieved as much as possible, and the feasibility of imputation can be guaranteed. The advantages and disadvantages of different sub-imputation models in terms of precision and usability can be complementary. After all the sub-models have attempted to impute the missing values to be imputed, there may still be some missing values that all sub-imputation models fail to impute. In this case, a backup imputation model (i.e., other imputation models) can be used for imputation.

[0221] The above introduced the use of sub-imputation models in the integrated imputation model to process missing values. The following introduces how to construct an integrated imputation model.

[0222] In the embodiments of the present application, take the sub-imputation model including a regression imputation model and a time series imputation model as an example. The regression imputation model includes a LightGBM model, a K-nearest neighbor model, and a random forest. The time series imputation model includes an adjacent frame linear interpolation model and a cubic Lagrange interpolation model.

[0223] It should be noted that the time series imputation model can be directly used for missing value imputation, while the regression imputation model needs to be trained before it can be used. Therefore, the model training mentioned below is mainly for the regression imputation model, and the time series imputation model can directly calculate the imputed value corresponding to the missing value based on a preset formula. For example, the adjacent frame linear interpolation model satisfies the following formula 5, and the cubic Lagrange interpolation model satisfies the following formula 6.

[0224] (Formula 5).

[0225] Among them, is the acquisition time of the data frame where the missing value is located, is the acquisition time of the previous data frame of the data frame where the missing value is located, is the acquisition time of the next data frame of the data frame where the missing value is located, ; is the value of the parameter corresponding to the missing value in the previous data frame, is the value of the parameter corresponding to the missing value in the next data frame, is the imputed value corresponding to the missing value.

[0226] (Formula 6).

[0227] Among them, is the acquisition time of the th data frame before the data frame where the missing value is located, is the acquisition time of the The acquisition time of a data frame is the value of the parameter corresponding to the missing value in the th data frame.

[0228] In some embodiments, each data frame contains multiple parameters, and each type of parameter corresponds to an integrated imputation model. Taking the construction of the target integrated imputation model corresponding to the target parameter as an example, as Figure 8 shown, the target integrated imputation model can be obtained through the following S801 - S806.

[0229] S801. Determine the training data frames and test data frames from multiple data frames collected at different times.

[0230] Among them, the training data frames and test data frames contain multiple data frames in which the target parameter is not a missing value. The training set is used to train the sub - imputation model, and the test set is used to determine the accuracy index of the trained sub - imputation model.

[0231] Specifically, the server can use the random sampling without replacement or random sampling with replacement method to obtain the sampled data frames from multiple data frames collected at different times. Then, the server can divide the sampled data frames into training data frames and test data frames based on a preset division ratio. Then, for each data frame in the test data frames, the server can obtain the acquisition time of the data frame, as well as the acquisition times of the two data frames before and after the data frame among the multiple data frames.

[0232] The embodiments of the present application do not limit the preset division ratio. For example, the sampled data frames can be divided according to a ratio of 5:1, or can be divided according to a ratio of 4:1.

[0233] In the embodiments of the present application, it is possible to determine whether to use the random sampling without replacement or random sampling with replacement method based on the number of multiple data frames and the number of vehicles to which the multiple data frames belong. For example, if the ratio of the number of multiple data frames to the number of vehicles is greater than or equal to the quantity threshold, then use random sampling without replacement; if the ratio of the number of multiple data frames to the number of vehicles is less than the quantity threshold, then use random sampling with replacement.

[0234] In the embodiments of the present application, before training the sub - imputation model, it is also necessary to perform standardization processing on other parameters in the training data frames except the target parameter. Taking the K - nearest neighbor model as an example. Before training the K - nearest neighbor model corresponding to the target parameter, the server can first perform standardization processing on other parameters except the target parameter, and the standardization calculation formula can refer to the following formula 7.

[0235] (Formula 7).

[0236] Among them, is a parameter among multiple parameters, is the corresponding parameter mean in the training data frame, is the corresponding parameter standard deviation in the training data frame, is the parameter after standardization.

[0237] S802. Screen out the optimal feature parameters from other parameters included in the training data frame except the target parameter, and use them as the input features of the sub-imputation model.

[0238] Specifically, the server can use a preset feature selection method to screen out the optimal features from each parameter included in the training data frame as the input features of the sub-imputation model, so as to delete redundant features and ensure the accuracy and efficiency of the sub-imputation model.

[0239] In the embodiments of the present application, the preset feature selection method may include methods such as manual feature selection and backward sequential feature elimination, which are not limited herein. For example, the method of using the backward sequential feature elimination method to determine the optimal features can refer to the following embodiments shown in Figure 9 and will not be elaborated herein.

[0240] S803. For each original sub-imputation model, use the optimal feature parameters in the training data frame to train the original sub-imputation model to obtain the trained sub-imputation model.

[0241] Specifically, the server can input the optimal feature parameters in the training data frame into each original imputation model to obtain each trained sub-imputation model.

[0242] S804. For each trained sub-imputation model, input the optimal feature parameters in the test data frame into the trained sub-imputation model to obtain the imputation result corresponding to the target parameter.

[0243] S805. For each trained sub-imputation model, based on the target parameter included in the test data frame and the imputation result corresponding to the target parameter, determine the accuracy index of the trained sub-imputation model.

[0244] Exemplarily, taking the accuracy index as the root mean square error RMSE, the smaller the RMSE, the higher the model accuracy. Table 3 shows the root mean square error RMSE of different sub-imputation models under different parameters.

[0245] Table 3 Root Mean Square Error RMSE of Different Sub-imputation Models under Different Parameters

[0246] S806. Combine in the order of the accuracy indexes of each trained sub-imputation model to obtain the target integrated imputation model corresponding to the target parameter.

[0247] Exemplarily, taking the driving motor speed as the target parameter, in combination with Table 3 above. Since the smaller the RMSE, the higher the model accuracy, therefore, as Figure 10 shown, the order of the sub-interpolation models included in the target integrated interpolation model is Random Forest, K-Nearest Neighbor model, LightGBM model, adjacent frame linear interpolation model, and cubic Lagrange interpolation model. In addition, the wide-range linear interpolation model and the nearest neighbor interpolation model can also be used as backup models and integrated into the target integrated interpolation model.

[0248] As Figure 10 shown, in the case where the target parameter in data frame A is a missing value, the server can, in accordance with the order of the integrated interpolation model, successively use the Random Forest, K-Nearest Neighbor model, LightGBM model, adjacent frame linear interpolation model, cubic Lagrange interpolation model, wide-range linear interpolation model, and nearest neighbor interpolation model to interpolate the missing value until the interpolation result corresponding to the target parameter is obtained. Among them, the parameters input to each sub-interpolation model include other parameters in data frame A except the target parameter, and the target parameters included in the previous data frame B and the subsequent data frame C adjacent to data frame A.

[0249] Based on the above technical means, the present application can solve the situation where the parameters included in the data frame have missing values by constructing an integrated interpolation model. Compared with directly deleting the data frame containing the missing value, using the interpolation model for processing can retain more valid data and reduce data waste.

[0250] In some embodiments, taking the backward sequential feature elimination method as an example, as Figure 9 shown, using the backward sequential feature elimination method to obtain the optimal features may specifically include: Step 1 - Step 10.

[0251] Step 1: Obtain the sub-training data frame for optimal feature selection from the training data frame.

[0252] It should be noted that the feature selection method based on backward sequential feature elimination has a large computational workload. Therefore, the server can randomly sample from the training set to obtain the sub-training data frame for optimal feature selection. Among them, the sample size of the sub-training data frame can be determined according to the actual computing resources and is not limited thereto.

[0253] Step 2: Use the sub-training data frame to train the sub-interpolation model and determine the initial accuracy index of the sub-interpolation model.

[0254] In the embodiments of the present application, the accuracy index may include root mean squared error (RMSE), mean squared error (MSE), etc., and is not limited thereto.

[0255] Specifically, the server may input other parameters in the sub-training data frame except for the target parameter into the sub-imputation model to obtain the target parameter predicted by the model. Then, the server may calculate the accuracy index of the sub-imputation model based on the target parameter in the sub-training data frame and the target parameter predicted by the model.

[0256] Step 3: Use a preset hyperparameter optimization algorithm to optimize the hyperparameters of the sub-imputation model.

[0257] Among them, hyperparameters refer to the parameters set before training the model, which are used to control the behavior and performance of the model. The selection of hyperparameters can affect the training speed, convergence, capacity, and generalization ability of the model, etc.

[0258] Specifically, the server may first determine the target hyperparameters and the search range corresponding to each target hyperparameter. Then, the server uses a preset hyperparameter optimization algorithm to determine the optimal value of each hyperparameter from the search range corresponding to each target hyperparameter. Table 4 shows the target hyperparameters that need to be optimized for the K-nearest neighbor model, random forest, and LightGBM model.

[0259] Among them, the target hyperparameters corresponding to the K-nearest neighbor model may include the number of neighboring data (n_neighbors), weights (weights), and distance metric parameter (p). The target hyperparameters corresponding to the random forest may include the number of iterations (n_estimators), the maximum depth of the tree (max_depth), the minimum number of samples in the leaf node (min_samples_leaf), the maximum number of features (max_features), and the maximum number of samples in each tree (max_samples). The target hyperparameters corresponding to the LightGBM model may include n_estimators, the number of leaf nodes in each tree (num_leaves), max_depth, learning rate (learning_rate), the minimum number of samples required for the leaf node (min_child_samples), the sample ratio used when training each tree (subsample), and the subsample frequency, that is, how many rounds of iterations to perform a subsampling (subsample_freq).

[0260] Table 4 Target hyperparameters that need to be optimized for the K-nearest neighbor model, random forest, and LightGBM model

[0261] In the embodiments of the present application, the preset hyperparameter optimization algorithm may include a random search algorithm, a Bayesian optimization algorithm, a grid search algorithm, etc., which are not limited thereto.

[0262] Exemplarily, taking the sub-interpolation model as the LightGBM model and the preset hyperparameter optimization algorithm as the Bayesian optimization algorithm as an example. The server can use the Python language to call the lightgbm package to implement the LightGBM model. Then the server can call the Bayesian optimization algorithm (tree-structured parzen estimator, TPE) provided by the optuna package to search for the optimal hyperparameters of the LightGBM model. Among them, in the process of searching for hyperparameters, the accuracy index of each group of hyperparameters can be obtained by using the K-fold cross-validation method, and the accuracy index can be the root mean square error RMSE.

[0263] Step 4. Initialize the parameter sequence i = 1.

[0264] Step 5. Determine whether i is greater than the remaining parameters i _ max in the sub-training data frame. If so, execute Step 8; if not, execute Step 6.

[0265] Step 6. Use the other parameters in the sub-training data frame except the i-th parameter to train the sub-interpolation model, and determine the accuracy index corresponding to the sub-interpolation model.

[0266] Step 7. Let i = i + 1.

[0267] Step 8. Determine the redundant parameters so that the sub-interpolation model trained based on the other parameters except the redundant parameters has the highest accuracy index, and delete the redundant parameters in the sub-training data frame.

[0268] Step 9. Determine the remaining parameter i _ max = 1 or whether the consecutive decrease times of the accuracy index of the sub-interpolation model are greater than the times threshold. If so, execute Step 10; if not, execute Step 3.

[0269] Among them, the times threshold can be 5 times, 6 times, etc., and this application does not make a limitation on this.

[0270] Step 10. Take the input features corresponding to the sub-interpolation model with the highest accuracy index as the optimal parameter features of the sub-training data frame.

[0271] Combining the above Steps 1 - 10, it can be understood that backward sequential feature elimination can be performed in multiple rounds of iteration. By deleting one parameter in the sub-training in each round of iteration, the original parameters in the sub-training data frame are refined into the optimal parameter features to improve the accuracy of the sub-interpolation model. The finally obtained optimal parameter features are the input features corresponding to the model with the highest accuracy index in all iterations.

[0272] In the embodiments of the present application, the optimal feature parameters (i.e., the model input features) may vary depending on the training data frames and the regression imputation model. Table 5 shows the feature selection results of the regression imputation sub-model. In Table 5, the driving motor speed -1 is the driving motor speed of the data frame before the target data frame, and the driving motor speed +1 is the driving motor speed of the data frame after the target data frame. The same applies to other cases and will not be elaborated here. The input feature is the optimal feature, and the target parameter is the output parameter of the model, that is, the target parameter predicted by the model can be obtained after inputting the input feature into the model.

[0273] Table 5 Feature Selection Results of the Regression Imputation Sub-Model

[0274] Continued Table 5

[0275] Based on the above technical solution, the present application can delete redundant feature parameters and use the optimal feature parameters as input features to train the model, so as to improve the processing efficiency and output accuracy of the sub-imputation model.

[0276] In an alternative embodiment, before performing the above S201, the data processing method provided by the present application further includes: the server can decode each original data frame according to a preset decoding rule to obtain a plurality of decoded data frames, and each decoded data frame corresponds to a vehicle identification code. Then, the server can store the data frames with the same vehicle identification code correspondingly to obtain the data frames corresponding to each vehicle. On this basis, the above S201 may include: for the target vehicle among the multiple vehicles, the server can obtain the first data frame from the data frames corresponding to the target vehicle.

[0277] In some embodiments, the server stores the original data frames collected by different vehicles at different times, and each original data frame is an encoded data frame. Based on this, the server can decode each original data frame according to a preset decoding rule. Then, the server can segment the decoded data frames according to the vehicle identification code (vendor identification, VID) in each decoded data frame to obtain the data frames corresponding to each vehicle. Then, for each vehicle, the server can store them structurally according to the collection time of the data frames corresponding to the vehicle.

[0278] Exemplarily, Table 6 shows the encoding rules for each parameter in the original data frame based on the GB / T 32960 standard. The server can decode the parameters in each data frame (such as scaling and offset) based on the encoding rules to obtain the true values of each parameter in common units. Combining the following Table 6, taking the driving motor speed as an example, the effective value range of the encoded driving motor speed is between 0 and 65531, and its data offset is 20000. The server can subtract the data offset of 20000 from the encoded driving motor speed to obtain the decoded driving motor speed, whose range is between -20000 r / min and 45531 r / min. After that, the server can store the decoded data in the structured data table shown in Table 7. Each row in the structured data table is a data frame, and each data frame contains various parameters of the vehicle at a certain acquisition moment.

[0279] Table 6 Encoding Rules for Each Parameter in the Original Data Frame Based on the GB / T 32960 Standard

[0280] Continued Table 6

[0281] Table 7 Structured Data Table

[0282] Based on the above technical solution, by decoding the original data frame according to the preset decoding rules and storing the data frame corresponding to the vehicle identification code, the effective classification of vehicle data is realized, which is convenient for subsequent data management and analysis of specific vehicles. For the target vehicle among multiple vehicles, the first data frame can be obtained from its corresponding data frame. This method makes the processing of specific vehicle data more efficient and accurate, and improves the data processing efficiency.

[0283] The above respectively introduced the processing methods of data frames with the same acquisition time (and duplicate data), the processing methods of missing values for the parameters included in the data frame, the driving segment division method, and the method of obtaining the data frame corresponding to each vehicle. The following will introduce the data processing method provided by this application in combination with the above various embodiments. As Figure 11 shown, the data processing method provided by this application mainly includes data structuring processing (S1101 - S1102), driving segment division (S1103 - S1104), outlier processing (S1105 - S1106)), and missing value imputation (S1107 - S1108). This method can specifically include: S1101. Decode the original data according to the preset decoding rules.

[0284] S1102. Integrate the decoded data into a structured data table with the acquisition time series as rows and various parameters as columns. Each row in the structured data table is a data frame.

[0285] S1103. Correct the vehicle state and charging state corresponding to each data frame, and determine the vehicle running state corresponding to each data frame based on the vehicle state and charging state corresponding to each data frame.

[0286] S1104. For any two adjacent data frames among multiple data frames, determine whether the two adjacent data frames belong to the same driving segment based on the vehicle running states corresponding to the two adjacent data frames and the time interval between the two adjacent data frames.

[0287] S1105. For each data frame in the same driving segment, determine the parameters exceeding the threshold range in the data frame as outliers, and replace the parameter with a missing value.

[0288] S1106. Obtain the first data frame from the same driving segment, and determine the second data frame based on the adjacent data frame and the first data frame of the first data frame, and retain the data frame with the highest similarity to the second data frame in the first data frame.

[0289] S1107. Construct an integrated imputation model.

[0290] S1108. For each data frame in the same driving segment, when the parameter included in the data frame is a missing value, use the integrated imputation model to process the missing value.

[0291] Figure 12 The structural schematic diagram of a data processing device provided by an embodiment of the present application is as Figure 12 shown. The device includes: an acquisition unit 1201, a first determination unit 1202, a second determination unit 1203, and a processing unit 1204.

[0292] The acquisition unit 1201 is configured to acquire a first data frame, and the first data frame includes multiple data frames collected by the target vehicle at the first moment.

[0293] The first determination unit 1202 is configured to determine a second data frame based on the adjacent data frame and the first data frame of the first data frame. The adjacent data frame of the first data frame includes the data frames collected by the target vehicle at the adjacent moment of the first moment.

[0294] The second determination unit 1203 is configured to determine a target data frame with the highest similarity to the second data frame from the first data frame.

[0295] The processing unit 1204 is configured to use the target data frame as the data frame at the first moment.

[0296] In a possible way, the first determination unit 1202 is further configured to, when the parameters included in multiple data frames satisfy the first judgment condition, determine a second data frame based on the adjacent data frames of the first data frame and the first data frame.

[0297] In a possible way, the first determination unit 1202 is further configured to, when the number of adjacent data frames of the first data frame is less than a threshold, generate a third data frame based on the adjacent data frames of the first data frame. Then, determine the second data frame based on the adjacent data frames of the first data frame, the first data frame, and the third data frame.

[0298] In a possible way, the second determination unit 1203 is further configured to, for each of the data frames in the first data frame, determine the degree of difference between the data frame and the second data frame, and use the data frame with the smallest degree of difference between the first data frame and the second data frame as the target data frame.

[0299] In a possible way, the acquisition unit 1201 is further configured to acquire multiple data frames collected by the target vehicle at different times.

[0300] In a possible way, the first determination unit 1202 is further configured to, for each of the multiple data frames collected by the target vehicle at different times, determine the vehicle running state corresponding to the data frame based on the vehicle state and the charging state corresponding to the data frame.

[0301] In a possible way, the first determination unit 1202 is further configured to, for any two adjacent data frames among the multiple data frames collected by the target vehicle at different times, determine whether the two adjacent data frames belong to the same driving segment based on the vehicle running states corresponding to the two adjacent data frames and the time interval between the two adjacent data frames.

[0302] In a possible way, the acquisition unit 1201 is further configured to acquire the first data frame from the multiple data frames included in the same driving segment.

[0303] In a possible way, the processing unit 1204 is further configured to, for each of the data frames included in the same driving segment, when the vehicle state and / or the charging state corresponding to each data frame satisfy the preset correction condition, correct the vehicle state and / or the charging state corresponding to each data frame based on the charging state, the vehicle speed, and the total current corresponding to each data frame.

[0304] In a possible way, the processing unit 1204 is further configured to, for each of the data frames included in the first data frame, when the parameter included in the data frame is a missing value, successively use the sub-imputation models included in the integrated imputation model to perform imputation processing on the missing value until an imputation result is obtained.

[0305] Figure 13 This is a block diagram of an electronic device provided by an embodiment of the present application. As Figure 13 shown, the electronic device includes, but is not limited to: a processor 1301 and a memory 1302.

[0306] Among them, the above-mentioned memory 1302 is used to store executable instructions of the above-mentioned processor 1301. It can be understood that the above-mentioned processor 1301 is configured to execute instructions to implement the vehicle control method in the above-mentioned embodiment.

[0307] It should be noted that those skilled in the art can understand that Figure 13 the structure of the electronic device shown in Figure 13 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than

[0308] shown, or combine some components, or have different component arrangements.

[0309] The processor 1301 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1302, and calling data stored in the memory 1302, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 1301 may include one or more processing units. Optionally, the processor 1301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1301 either.

[0310] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as the memory 1302 including instructions. The above-mentioned instructions can be executed by the processor 1301 of the electronic device to implement the method in the above-mentioned embodiment.

[0311] In actual implementation, Figure 12 the functions of the acquisition unit 1201, the first determination unit 1202, the second determination unit 1203, and the processing unit 1204 inFigure 13 The processor 1301 in Figure 13 calls the computer program stored in the memory 1302 to implement. For the specific execution process, reference can be made to the description in the method section of the above embodiment, which will not be elaborated here.

[0312] Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0313] In an exemplary embodiment, the embodiment of the present application also provides a computer program product including one or more instructions, and the one or more instructions can be executed by the processor 1301 of the electronic device to complete the method in the above embodiment.

[0314] It should be noted that when the instructions in the above computer-readable storage medium or the one or more instructions in the computer program product are executed by the processor of the electronic device, each process of the above method embodiment is implemented, and the same technical effects as the above method can be achieved. To avoid repetition, it will not be elaborated here.

[0315] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0316] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0317] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0318] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0319] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0320] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire a first data frame, wherein the first data frame includes a plurality of data frames collected by the target vehicle at a first moment; Determine a second data frame based on adjacent data frames of the first data frame and the first data frame; the adjacent data frames of the first data frame include data frames collected by the target vehicle at adjacent moments of the first moment; Determine, from the first data frame, a target data frame having the highest similarity to the second data frame; The target data frame is used as the data frame at the first moment.

2. The method according to claim 1, characterized in that: The determining of the second data frame based on the adjacent data frame of the first data frame and the first data frame includes: When the parameters included in the multiple data frames meet the first judgment condition, the second data frame is determined based on the adjacent data frames of the first data frame and the first data frame, and each of the data frames includes continuous parameters and / or discrete parameters.

3. The method according to claim 2, characterized in that The first judgment condition includes: The range of at least one continuous parameter included in the multiple data frames is greater than the range threshold; and / or, The multiple data frames contain different discrete parameters.

4. The method according to claim 2 or 3, characterized in that: The method further comprises: When the parameters included in the multiple data frames meet the second judgment condition, any data frame is selected from the first data frames as the data frame at the first moment.

5. The method according to claim 4, characterized in that Each of the data frames contains continuous parameters and / or discrete parameters, and the second judgment condition includes: The range of each continuous parameter included in the multiple data frames is less than or equal to the range threshold, and each discrete parameter included in the multiple data frames is the same.

6. The method according to claim 1, characterized in that The method further comprises: When the number of adjacent data frames of the first data frame is less than a threshold, generating a third data frame based on the adjacent data frames of the first data frame; The determining of the second data frame based on the adjacent data frame of the first data frame and the first data frame includes: The second data frame is determined based on adjacent data frames of the first data frame, the first data frame, and the third data frame.

7. The method according to any one of claims 1 to 3, characterized in that The determining, from the first data frame, a target data frame having the highest similarity to the second data frame comprises: For each of the data frames in the first data frames, determining a degree of difference between the data frame and the second data frame; the degree of difference is negatively correlated with the similarity; The data frame with the smallest difference between the first data frame and the second data frame is used as the target data frame.

8. The method according to claim 1, characterized in that: The method further comprises: Acquire multiple data frames collected from the target vehicle at different times; For each data frame of the plurality of data frames collected at different times by the target vehicle, based on the vehicle state and charging state corresponding to the data frame, determine the vehicle operation state corresponding to the data frame; the vehicle state refers to the state of the vehicle in different working modes; the vehicle operation state includes different states of the vehicle during operation; For any two adjacent data frames among a plurality of data frames collected by the target vehicle at different times, based on the vehicle running states corresponding to the any two adjacent data frames and the time interval between the any two adjacent data frames, determine data frames belonging to the same driving segment; The obtaining of the first data frame comprises: The first data frame is acquired from a plurality of data frames included in the same driving segment.

9. The method according to claim 8, characterized in that The vehicle operating state includes a first operating state and a second operating state. The first operating state means that the vehicle state is in a started state, and the charging state is a driving charging state or an uncharged state. The driving charging state means that the battery is in an energy recovery state during the vehicle driving process; the second operating state means that the vehicle state is in an unstarted state.

10. The method according to claim 8 or 9, characterized in that: Before determining the vehicle running state corresponding to the data frame based on the vehicle state and charging state corresponding to the data frame, the method further includes: For each data frame contained in the same driving segment, when the vehicle state and / or charging state corresponding to each data frame meets the preset correction condition, the vehicle state and / or charging state corresponding to each data frame is corrected based on the charging state, vehicle speed and total current corresponding to each data frame.

11. The method according to claim 10, characterized in that The modification conditions include any of the following: The vehicle state corresponding to each of the data frames contained in the same driving segment is other vehicle states, and the other vehicle states refer to vehicle states other than the start state and the shutdown state; The vehicle state corresponding to each of the data frames included in the same driving segment is a start state and the charging state is another charging state, and the another charging state refers to a charging state other than a driving charging state and an uncharged state.

12. The method according to claim 1, characterized in that The method further comprises: For each of the data frames included in the first data frame, when a parameter included in the data frame is a missing value, interpolation processing is performed on the missing value by using the sub-interpolation model included in the integrated interpolation model one by one until an interpolation result is obtained, wherein the missing value includes a parameter whose value is empty or a parameter whose value exceeds a preset range; The order of the sub-interpolation models for interpolating the missing values ​​is related to the accuracy index of the sub-interpolation models, and the accuracy index is used to characterize the accuracy of the interpolation results obtained by using the sub-interpolation models.

13. A data processing device, characterized in that: The device comprises: an acquisition unit, a first determination unit, a second determination unit and a processing unit; The acquisition unit is used to acquire a first data frame, wherein the first data frame includes a plurality of data frames collected by the target vehicle at a first moment; The first determining unit is used to determine a second data frame based on an adjacent data frame of the first data frame and the first data frame; the adjacent data frame of the first data frame includes a data frame collected by the target vehicle at an adjacent time to the first time; The second determining unit is used to determine, from the first data frame, a target data frame having the highest similarity to the second data frame; The processing unit is used to use the target data frame as the data frame at the first moment.

14. An electronic device, characterized in that: including memory and processor; The memory is coupled to the processor; The memory is used to store computer program code, wherein the computer program code includes computer instructions; When the processor executes the computer instructions, the electronic device performs the data processing method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that: When the computer-executable instructions stored in the computer-readable storage medium are executed by a processor of a processing device, the processing device can execute the data processing method according to any one of claims 1 to 12.

16. A computer program product, characterized in that The computer program product comprises the computer program, and the computer program is suitable for being loaded by a processor and executing the data processing method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image processing unit, image pickup device, image processing program, and image processing method

    CN106852190A

  • Monocular vision odometer pose processing method based on IMU assistance

    CN110009681A

  • Video compression method with low loss

    CN111901600A

  • Joint labeling method and device for vehicle-road coordination data

    CN114511765A

  • Multimedia data optimization processing method and device for network set top box

    CN115842909A