A method and apparatus for processing vehicle data

CN116089693BActive Publication Date: 2026-09-11SHANGHAI QINGGAN INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111287411.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-02
Publication Date
2026-09-11
Estimated Expiration
2041-11-02

AI Technical Summary

Technical Problem

这种现有的处理方式不但会降低数据数量及质量、浪费脏数据中仍存在的有价值信息,还会直接影响着推荐系统的推荐准确性及有效性

Benefits of technology

[0009] The computer-readable storage medium provided in the third aspect of the present invention stores computer instructions thereon. When the computer instructions are executed by a processor, the vehicle data processing method provided in the first aspect of the present invention is implemented. By implementing the above processing method, the computer-readable storage medium can optimize the raw data obtained by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors, thereby improving the quantity and quality of data and making full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116089693B_ABST
    Figure CN116089693B_ABST
Patent Text Reader

Abstract

The application provides a vehicle data processing method and device, and a computer readable storage medium. The processing method comprises the following steps: obtaining first original data of a first vehicle; determining whether dirty data exists in the first original data; in response to a determination result that dirty data exists in the first original data, performing cluster analysis on standard data in the first original data and a plurality of groups of second original data of at least one second vehicle and / or a plurality of groups of third original data of the first vehicle according to a data dimension of the dirty data, determining at least one group of cluster data of the first original data from the plurality of groups of second original data and / or the plurality of groups of third original data; and optimizing the first original data according to the at least one group of cluster data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to vehicle data processing technology, and more particularly to a vehicle data processing method, a vehicle data processing apparatus, and a computer-readable storage medium. Background Technology

[0002] The Internet of Vehicles (IoV) is a network of interconnected vehicles that uses moving vehicles as the information sensing objects and leverages next-generation information and communication technologies to achieve network connectivity between vehicles, between vehicles and people, between vehicles and roads, and between vehicles and service platforms. It can improve the overall intelligent driving level of vehicles, provide users with a safe, comfortable, intelligent, and efficient driving experience and transportation services, while improving traffic operation efficiency and enhancing the intelligence level of social transportation services.

[0003] Intelligent recommendation systems based on the Internet of Vehicles (IoV) can acquire a large amount of vehicle and user data from the IoV and use pre-trained artificial intelligence models to provide intelligent vehicle recommendations based on vehicle and user data collected by the vehicle's sensors. However, due to network, storage, processing, and data acquisition equipment failures, the raw data acquired by existing intelligent recommendation systems from the IoV and vehicle sensors often contains a large amount of "dirty data," including missing fields, missing records, data anomalies, and data contradictions. Currently, the only approach to handling this dirty data is to discard or ignore it. This existing method not only reduces the quantity and quality of data and wastes valuable information still present in the dirty data, but also directly affects the accuracy and effectiveness of the recommendation system.

[0004] In order to overcome the above-mentioned defects in the existing technology, there is an urgent need in the field for a vehicle data processing technology to optimize the raw data obtained by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors, so as to improve the quantity and quality of data and make full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system. Summary of the Invention

[0005] The following provides a brief overview of one or more aspects to offer a basic understanding of them. This overview is not an exhaustive summary of all conceived aspects, nor is it intended to identify key or decisive elements of all aspects, nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.

[0006] To overcome the aforementioned deficiencies in the existing technology, the present invention provides a vehicle data processing method, a vehicle data processing device, and a computer-readable storage medium, which can optimize the raw data obtained by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors to improve the quantity and quality of data, and make full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system.

[0007] Specifically, the vehicle data processing method provided by the first aspect of the present invention includes the following steps: acquiring first raw data of a first vehicle; determining whether dirty data exists in the first raw data; in response to the determination result that dirty data exists in the first raw data, performing cluster analysis on the normalized data in the first raw data with multiple sets of second raw data of at least one second vehicle and / or multiple sets of third raw data of the first vehicle according to the data dimension of the dirty data, determining at least one set of co-clustered data of the first raw data from the multiple sets of second raw data and / or the multiple sets of third raw data; and optimizing the first raw data according to the at least one set of co-clustered data. By performing these steps, the vehicle data processing method can optimize the raw data acquired by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors, thereby improving the quantity and quality of data and making full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system.

[0008] The vehicle data processing apparatus provided in the second aspect of the present invention includes a memory and a processor. The processor is connected to the memory and configured to implement the vehicle data processing method provided in the first aspect of the present invention. By implementing the processing method, the processing apparatus can optimize the raw data obtained by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors, thereby improving the quantity and quality of data and making full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system.

[0009] The computer-readable storage medium provided in the third aspect of the present invention stores computer instructions thereon. When the computer instructions are executed by a processor, the vehicle data processing method provided in the first aspect of the present invention is implemented. By implementing the above processing method, the computer-readable storage medium can optimize the raw data obtained by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors, thereby improving the quantity and quality of data and making full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system. Attached Figure Description

[0010] The above-described features and advantages of the present invention will be better understood after reading the following detailed description of embodiments of the present disclosure in conjunction with the accompanying drawings. In the drawings, components are not necessarily drawn to scale, and components having similar related characteristics or features may have the same or similar reference numerals.

[0011] Figure 1 A flowchart illustrating the training phase of a vehicle data processing method provided according to some embodiments of the present invention is shown.

[0012] Figure 2 A flowchart illustrating the optimization phase of a vehicle data processing method provided according to some embodiments of the present invention is shown. Detailed Implementation

[0013] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention is presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to these embodiments. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a thorough understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description.

[0014] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0015] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood as the orientations shown in the relevant paragraphs and accompanying drawings. These relative terms are for illustrative purposes only and do not imply that the described apparatus must be manufactured or operated in a specific orientation, and therefore should not be construed as limiting the invention.

[0016] It is understood that although terms such as "first," "second," and "third" may be used herein to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first components, regions, layers, and / or parts discussed below may be referred to as second components, regions, layers, and / or parts without departing from some embodiments of the present invention.

[0017] As mentioned above, existing dirty data processing techniques in this field can only discard or ignore dirty data. This existing approach not only reduces the quantity and quality of data and wastes valuable information still present in dirty data, but also directly affects the accuracy and effectiveness of recommendation systems.

[0018] To overcome the aforementioned deficiencies in the existing technology, the present invention provides a vehicle data processing method, a vehicle data processing device, and a computer-readable storage medium, which can optimize the raw data obtained by the intelligent recommendation system from the Internet of Vehicles and vehicle sensors to improve the quantity and quality of data, and make full use of the valuable information still existing in the dirty data, thereby improving the recommendation accuracy and effectiveness of the recommendation system.

[0019] In some non-limiting embodiments, the vehicle data processing method provided in the first aspect of the present invention can be implemented by the vehicle data processing apparatus provided in the second aspect of the present invention. Specifically, the vehicle data processing apparatus is configured with a dedicated or shared memory and a processor. The memory includes, but is not limited to, the computer-readable storage medium provided in the third aspect of the present invention, on which computer instructions are stored. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the vehicle data processing method provided in the first aspect of the present invention.

[0020] In some embodiments of the present invention, the vehicle data processing method provided in the first aspect of the present invention can be implemented in two parts: an offline training phase and an online optimization phase. Correspondingly, the vehicle data processing apparatus provided in the second aspect of the present invention can also be divided into a training apparatus and an optimization apparatus to implement the steps of the training phase and the optimization phase, respectively.

[0021] The working principle of the training device will first be described using some embodiments of the training phase of the vehicle data processing method. Those skilled in the art will understand that these vehicle data processing methods are merely non-limiting embodiments provided by this invention, intended to clearly demonstrate the main concepts of the invention and provide specific solutions convenient for public implementation, rather than limiting all functions or operating methods of the training device. Similarly, this training device is also only a non-limiting embodiment provided by this invention and does not limit the entities executing the steps in the above-described vehicle data processing method.

[0022] Please refer to Figure 1 , Figure 1 A flowchart illustrating the training phase of a vehicle data processing method provided according to some embodiments of the present invention is shown.

[0023] like Figure 1 As shown, in the training phase of the vehicle data processing method, the trainer can first create an initial quality optimization model via a training device. In some embodiments, this training device can be a computer device such as a personal computer (PC) or a workstation. The quality optimization model can be a clustering analysis algorithm model such as DBscan or Kmeans.

[0024] It is understood that the above-described trainer is only a non-limiting description provided by the present invention, and may include a technician performing the above-described training steps, or a processor and / or controller executing the corresponding training instructions.

[0025] After creating the initial quality optimization model, the training device will acquire a large number of raw data samples from sample vehicles through local databases, cloud databases of the vehicle network, and other means for training the aforementioned quality optimization model. Specifically, each set of raw data samples D i The data can contain information from multiple dimensions, including but not limited to the interior temperature information T obtained from the temperature sensor of vehicle i. i Time information t obtained from the timer of vehicle i i Vehicle speed information v obtained from the speed sensor of vehicle i i Mileage information obtained from the odometer sensor of vehicle i. i The remaining fuel level information g obtained from the fuel level sensor of vehicle i i Longitude information EW obtained from the positioning module of vehicle i i and latitude information SN i City information obtained from the map software of the vehicle's infotainment system (i). iDriver image data p acquired from the image acquisition module of vehicle i i One or more of the vehicle data, and the air conditioning switch command O provided by user i to the corresponding vehicle i. aci Temperature adjustment command O ti Seat adjustment command O si Acceleration command O ai and / or braking command O bi One or more of the user command data, and user i's action data a i and / or e-expression data i This includes one or more of the user behavior data. The training device can then arrange these pieces of information in sequence to form a multidimensional feature vector of a set of original data samples, such as D. i =(T i ,t i ,v i ,o i ,g i EW i ,SN i ,c i ,p i O aci O ti O si O ai O bi ,a i ,e i ).

[0026] After acquiring a large number of raw data samples from vehicles, the training device will adjust the learning parameters of the quality optimization model according to each group of raw data samples i to obtain the silhouette coefficients of the algorithm evaluation under each learning parameter. Then, based on the silhouette coefficients, a hyperparameter curve will be plotted to determine the learning parameters of the optimal solution, thereby training the quality optimization model to realize the function of cluster analysis of input data.

[0027] Taking the DBscan algorithm model as an example, the trainer can first set the range of hyperparameters such as the cluster radius and the minimum number of points within the radius using the training device. Then, the algorithm model is trained according to different learning parameter values ​​to obtain the silhouette coefficients for evaluating the algorithm model under each learning parameter value. Next, the training device will plot hyperparameter curves based on these silhouette coefficients to obtain the optimal solution for each learning parameter value. Then, the training device will filter and compare the optimal solutions according to preset error conditions to obtain the optimal model that meets the error conditions, and determine the values ​​of each learning parameter in the optimal model, thus completing the training process of the DBscan algorithm model. The trained DBscan algorithm model can divide the large amount of original data samples into multiple clusters based on the values ​​of the learning parameters determined during training, and perform cluster analysis on the input original data to determine at least one set of data in the same cluster.

[0028] Furthermore, after determining the values ​​of each learning parameter in the optimal model, the training device can also perform cluster analysis on each original data sample based on the learning parameters of the optimal solution to divide them into multiple clusters of data within the same cluster. Then, the training device can obtain feedback data from the DBscan algorithm model to determine the contribution of each data dimension in each original data sample of each cluster to the cluster analysis result, and identify the multiple data dimensions with the highest contribution (e.g., 2 to 20) as mutually correlated dimensions.

[0029] The working principle of the above-mentioned optimization device will be further described below with reference to some embodiments of the optimization stage of the vehicle data processing method. Those skilled in the art will understand that these vehicle data processing methods are merely some non-limiting embodiments provided by the present invention, intended to clearly demonstrate the main concept of the invention and provide some specific solutions convenient for public implementation, rather than limiting all functions or all operating methods of the optimization device. Similarly, the optimization device is also only one non-limiting embodiment provided by the present invention and does not limit the executing entity of each step in the above-mentioned vehicle data processing method.

[0030] Please refer to Figure 2 , Figure 2 A flowchart illustrating the optimization phase of a vehicle data processing method provided according to some embodiments of the present invention is shown.

[0031] like Figure 2As shown, during the optimization phase of the vehicle data processing method, the optimization device can continuously acquire the vehicle's first raw data D1 while the vehicle is powered on and started, and determine whether there is dirty data in the first raw data D1. In some embodiments, the optimization device can be configured in the vehicle's in-vehicle infotainment system, the cloud server of the vehicle network, and / or a mobile terminal such as a mobile phone connected to the in-vehicle infotainment system, in the form of software programs and / or hardware devices. The first raw data D1 can be compared with the raw data sample D used in the training phase described above. i They have the same data structure, namely D1 = (T1, t1, v1, o1, g1, EW1, SN1, c1, p1, O ac1 O t1 O s1 O a1 O b1 (a1, e1). This dirty data can include various types such as missing records, missing fields, data exceeding limits, and data contradictions.

[0032] In some embodiments, when determining whether there is dirty data in the first raw data D1, the optimization device can first determine whether there is a record missing in the first raw data based on preset parameters such as upload time, time interval, and triggering conditions.

[0033] For example, in response to a vehicle's ignition operation, its in-vehicle system should generate a corresponding ignition record. This ignition record contains multiple fields such as mileage, remaining fuel, longitude, latitude, and temperature. If the ignition record cannot be obtained from the in-vehicle system, the optimization device can determine that there is dirty data in the first raw data D1, and the dirty data type is record missing, and determine the data dimension of the dirty data based on the dimension of the missing record.

[0034] Conversely, if the ignition record can be obtained from the vehicle system, the optimization device can determine that there is no missing record in the first raw data D1, and then further analyze the content of each data dimension in the first raw data D1 to obtain at least one field data, and further determine whether there is dirty data in the first raw data D1 based on whether the at least one field data meets the preset field standard.

[0035] For example, if the temperature sensor does not upload the temperature field T1, or if the data value recording the specific temperature is empty or NULL, the optimization device can determine that there is dirty data in the temperature dimension in the first raw data D1, and that the dirty data type is a missing field.

[0036] For example, if the fuel level sensor uploads the remaining fuel level field g1, but the data value of the specific remaining fuel level recorded therein is 140, which exceeds the preset value range of 0 to 60, then the optimization device can determine that there is dirty data in the remaining fuel level dimension in the first raw data D1, and its dirty data type is data exceeding the limit.

[0037] For example, if the speed field data value uploaded by the speed sensor is 50, while the longitude information EW1 and latitude information SN1 uploaded by the positioning module do not change, the optimization device can determine that there is dirty data in the first original data D1 in the vehicle speed dimension, longitude dimension and latitude dimension, and its dirty data type is data contradiction.

[0038] Conversely, if the field data of each data dimension in the first original data D1 all meet the preset field standards, the optimization device can determine that there is no dirty data in the first original data D1.

[0039] like Figure 2 As shown, in response to the determination that dirty data exists in the first raw data D1, the optimization device can first acquire multiple sets of second raw data D2 uploaded by at least one second vehicle, and / or multiple sets of third raw data D3 uploaded by the first vehicle at other times. Then, the optimization device can perform cluster analysis on the normalized data of other data dimensions in the first raw data D1 (excluding the dimension containing dirty data) with the multiple sets of second raw data D2 and / or the multiple sets of third raw data D3, and determine at least one set of data in the same cluster as the first raw data D1. The first raw data D1 is then optimized based on this at least one set of data in the same cluster.

[0040] Specifically, for the dirty data with missing temperature fields mentioned above, the optimization device can first normalize the data dimensions other than the dimension containing the dirty data in the first original data D1, i.e., D1' = (t1, v1, o1, g1, EW1, SN1, c1, p1, O ac1 O t1 O s1 O a1 O b1 Input the pre-trained quality optimization model (a1, e1). The quality optimization model can determine the temperature dimension T as the data dimension of the dirty data in the first original data D1 based on the default data dimension in the input data D1'. Then, based on the relevant parameters determined through pre-training, the quality optimization model can determine that the time dimension t, longitude dimension EW, latitude dimension SN, and city dimension c are the time information t1, longitude information EW1, latitude dimension SN1, and city dimension c1, which have the highest correlation with the temperature dimension T, are the relevant data of the dirty data T1.

[0041] Subsequently, the optimization device can perform cluster analysis based on the aforementioned time dimension t, longitude dimension EW, latitude dimension SN, and city dimension c. Based on the learning parameters determined during training in the quality optimization model, it can determine that the first original data D1 belongs to the region specified in t. i ∈(t1–Δt,t1+Δt),EW i ∈(EW1–ΔEW,EW1+ΔEW),SN i ∈(SN1–ΔSN,SN1+ΔSN), c i =c1. This data cluster includes the first original data D1, at least one set of the aforementioned second original data D2, and / or at least one set of the aforementioned third original data D3. In this way, the optimization device can determine each set of second original data D2 and / or each set of third original data D3 in this data cluster as data in the same cluster as the first original data D1.

[0042] like Figure 2 As shown, after determining at least one set of clustered data for the first original data D1, the optimization device can obtain normalized data T with the same dimension from each set of clustered data based on the data dimension T of the dirty data. i Afterwards, the optimization device can calculate the standard data T. i The average value is used to replace the dirty temperature data T1 in the first original data D1, so as to optimize the first original data D1.

[0043] Those skilled in the art will understand that the above-described scheme for determining at least one set of clustered data of the first original data D1 based on relevant dimensions such as time dimension t, longitude dimension EW, latitude dimension SN, and city dimension c is merely a non-limiting implementation method provided by the present invention. It is intended to clearly demonstrate the main concept of the present invention and provide a specific solution that is easy for the public to implement, rather than to limit the scope of protection of the present invention.

[0044] Optionally, in other embodiments, for dirty data related to the driver's gender dimension, the optimization device can first determine the relevant data for each relevant dimension in the first original data D1 based on the relevant dimensions such as seat height, seat back angle, and vehicle speed determined during training. Then, based on these relevant data, it can determine at least one set of clustered data for the first original data D1. Subsequently, the optimization device can replace the dirty gender data in the first original data D1 with the majority values ​​of the driver's gender dimension in these clustered data to achieve the effect of optimizing the first original data D1.

[0045] Optionally, in other embodiments, for dirty data in the driver gender dimension, the optimization device can also determine the relevant data of the driver image in the first original data D1 based on the relevant dimensions of the driver image determined by training, and perform feature engineering dimensionality reduction processing on the driver image data. After that, the optimization device can determine at least one set of clustered data in the first original data D1 based on the dimensionality reduction feature data of the driver image, and replace the dirty gender data in the first original data D1 with the majority values ​​of the driver gender dimension in these clustered data, so as to achieve the effect of optimizing the first original data D1.

[0046] Those skilled in the art will understand that the above-described training quality optimization model for determining at least one relevant dimension of dirty data is merely a non-limiting implementation of this invention, intended to clearly demonstrate the main concept of the invention and provide a specific solution that is easy for the public to implement, rather than intended to limit the scope of protection of this invention.

[0047] Optionally, in other embodiments, the function of determining at least one related dimension of the dirty data dimension can also be implemented through a pre-constructed dimension relationship correspondence table. Specifically, this dimension relationship correspondence table can be divided into two columns, where the first column records the data dimension of the dirty data (e.g., temperature dimension T), and the second column records the related data dimensions of the dirty data dimension (e.g., time dimension t, longitude dimension EW, latitude dimension SN, and city dimension c). In this way, the optimization device can determine at least one related data dimension of the dirty data dimension by looking up the table, and then perform cluster analysis on the first original data D1', multiple sets of second original data D2, and / or multiple sets of third original data D3 on the at least one related data dimension to determine at least one set of data in the same cluster of the first original data D1.

[0048] Compared to the training quality optimization model mentioned above for determining clustered data, although this table lookup method can record fewer and simpler correspondences, it has the advantages of low initial investment and low technical threshold. Technical personnel can choose the appropriate method to achieve the desired effect based on their own economic costs and technical capabilities.

[0049] Preferably, in some embodiments of the present invention, when determining the co-cluster data of the first original data D1 based on the data dimension of the dirty data, the optimization device can further determine the data dimension of the dirty data, and perform preliminary screening on the multiple sets of third original data D3 of the first vehicle and the multiple sets of second original data D2 of each second vehicle based on the data dimension of the dirty data, thereby reducing the proportion of invalid data in the clustering analysis process and improving the speed and efficiency of clustering analysis.

[0050] For example, in response to the dirty data's data dimension belonging to dimensions unrelated to other vehicle data, such as mileage, remaining fuel, longitude, or latitude, the optimization device can filter out invalid data provided by each second vehicle and perform cluster analysis only on the normal data D1' of the first original data D1 other than the dimension containing the dirty data, as well as multiple sets of third original data D3 provided by the first vehicle at other times, in order to determine at least one set of co-clustered data of the first original data D1 from the multiple sets of third original data D3.

[0051] For example, in response to the dirty data's data dimension being a data dimension involving absolute sensor error, such as temperature, driver age, or driver gender, the optimization device can filter out erroneous data provided by the first vehicle at other times, and only perform cluster analysis on the normal data D1' of the first raw data D1 other than the dimension containing the dirty data, and multiple sets of second raw data D2 provided by each second vehicle, so as to determine at least one set of data in the same cluster of the first raw data D1 from the multiple sets of second raw data D2.

[0052] Thus, the present invention can further reduce the proportion of invalid and erroneous data in the cluster analysis process, thereby improving the speed and efficiency of cluster analysis and the accuracy of cluster analysis results.

[0053] Furthermore, in some embodiments of the present invention, after optimizing the first original data D1 based on at least one set of clustered data, the vehicle data processing device can also acquire the optimized first original data D1” and provide corresponding recommendation information to the user of the first vehicle based on the optimized first original data D1”. This recommendation information includes, but is not limited to, the opening interfaces of recommended in-vehicle applications such as navigation applications, music applications, radio applications, shopping applications, food ordering applications, and chat applications; the usage interfaces of recommended vehicle functions such as air conditioning functions, windshield wiper functions, headlight functions, seat adjustment functions, and volume adjustment functions; and the usage interfaces of recommended vehicle services such as roadside assistance services, emergency assistance services, chauffeur services, car wash services, and vehicle repair services.

[0054] Based on the above description, the present invention can perform cluster analysis on the first original data according to the data dimension of the dirty data, determine at least one set of co-clustered data of the first original data from multiple sets of second and / or third original data provided by other vehicles and / or the vehicle itself, and optimize the first original data according to the at least one set of co-clustered data, so that the optimized dirty data regains usability, so as to make full use of the valuable information in the dirty data and increase the amount of usable data in the original data.

[0055] Furthermore, compared to traditional vehicle-to-everything (V2X) service recommendation technologies, the technology that uses the first original data optimized by this invention for V2X service recommendation has advantages such as large data volume and high data quality, which can significantly improve the accuracy and effectiveness of recommendation information.

[0056] Although the methods described above are illustrated and depicted as a series of actions for the sake of simplicity, it should be understood and appreciated that these methods are not limited by the order of the actions, as some actions may occur in a different order and / or concurrently with other actions from the illustrations and descriptions herein or not illustrated and described herein but which may be understood by those skilled in the art, according to one or more embodiments.

[0057] Those skilled in the art will understand that information, signals, and data can be represented using any of a variety of different techniques and arts. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0058] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.

[0059] Although the processing device, training device, and optimization device described in the above embodiments can be implemented through a combination of software and hardware, it is understood that the processing device, training device, and optimization device can also be implemented individually in software or hardware. For hardware implementation, the processing device, training device, and optimization device can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic devices for performing the above functions, or selected combinations of the above devices. For software implementation, the processing device, training device, and optimization device can be implemented using independent software modules such as procedures and functions running on a general-purpose chip, each module performing one or more functions and operations described herein.

[0060] The various illustrative logic modules and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.

[0061] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing vehicle data, characterized in that, Includes the following steps: Obtain the first raw data of the first vehicle; Determine whether there is dirty data in the first raw data; In response to the judgment that dirty data exists in the first raw data, cluster analysis is performed on the normalized data of other data dimensions in the first raw data other than the dimension where the dirty data is located, and multiple sets of second raw data of at least one second vehicle and / or multiple sets of third raw data of the first vehicle, according to the data dimension of the dirty data. The data dimensions of the dirty data are input into a pre-trained quality optimization model, and at least one relevant dimension corresponding to the data dimension is determined based on the learning parameters determined during training. From the first raw data, obtain at least one piece of relevant data for at least one relevant dimension; Based on the at least one piece of relevant data and the learning parameters determined by training, the data cluster to which the first original data belongs is determined, wherein the data cluster includes the first original data, and at least one set of the second original data and / or at least one set of the third original data; Each group of the second original data and / or each group of the third original data in the data cluster is determined to be data in the same cluster as the first original data; and Based on the at least one set of clustered data, the first raw data is optimized, wherein the first raw data and the second raw data respectively include vehicle data collected by at least one detector of the corresponding vehicle, and / or behavioral data of at least one user on the corresponding vehicle.

2. The processing method as described in claim 1, wherein, The at least one detector includes a temperature sensor, a timer, a speed sensor, an odometer sensor, a fuel level sensor, a positioning module, and / or an image acquisition module. The vehicle data includes in-vehicle temperature information, vehicle speed information, mileage information, remaining fuel information, longitude information, latitude information, city information, and / or driver image data, and / or The behavioral data includes air conditioning switch commands, temperature adjustment commands, seat adjustment commands, acceleration commands and / or braking commands provided by the at least one user to the corresponding vehicle, and / or the action data and / or facial expression data of the at least one user.

3. The processing method as described in claim 1, wherein, The step of determining whether there is dirty data in the first original data includes: Determine whether there are any missing records in the first original data; and In response to the situation where the first original data has missing records, it is determined that there is dirty data in the first original data, and the data dimension of the dirty data is determined based on the missing records.

4. The processing method as described in claim 3, wherein, The step of determining whether there is dirty data in the first raw data further includes: In response to the case where the first original data does not contain the missing record, the first original data is parsed to obtain at least one field data therein; Determine whether the at least one field data conforms to a preset field standard; and In response to the judgment result that any of the field data does not conform to the field standard, it is determined that there is dirty data in the first original data, and the data dimension is determined according to the record to which the field data that does not conform to the field standard belongs.

5. The processing method as described in claim 1, wherein, The quality optimization model is the DBscan algorithm model, and the learning parameters include the cluster radius and the minimum number of points contained within the radius. The steps for training the quality optimization model include: Create an initial quality optimization model; Acquire multiple raw data samples from multiple sample vehicles, wherein the multiple raw data samples include vehicle data samples collected by at least one detector of the multiple sample vehicles, and / or behavioral data samples of at least one user on each of the sample vehicles. Based on the multiple data samples, the learning parameters of the quality optimization model are adjusted to obtain the silhouette coefficients of the algorithm evaluation under each learning parameter; and Hyperparameter curves are plotted based on the profile coefficients to determine the learning parameters of the optimal solution.

6. The processing method as described in claim 5, wherein, After performing the step of plotting hyperparameter curves based on the contour coefficients to determine the learning parameters of the optimal solution, the step of training the quality optimization model further includes: Cluster analysis is performed based on the learning parameters of the optimal solution to divide the multiple original data samples into multiple clusters. The contribution of each data dimension in each original data sample of each cluster to the clustering analysis results is obtained; and The data dimensions with the highest contribution were identified as relevant dimensions.

7. The processing method as described in claim 1, wherein, The step of optimizing the first original data based on the at least one set of clustered data includes: Based on the data dimensions of the dirty data, obtain the standardized data with the same dimensions from each group of clustered data; Calculate the average value of the standardized data with the same dimension for each of the above; and The dirty data in the first original data is replaced with the average value.

8. The processing method as described in claim 1, wherein, The step of performing cluster analysis on the normalized data in the first original data with multiple sets of second original data and / or multiple sets of third original data of the first vehicle based on the data dimension of the dirty data, and determining at least one set of co-clustered data of the first original data from the multiple sets of second original data and / or the multiple sets of third original data, includes: Determine the data dimension of the dirty data; In response to the determination that the data dimension of the dirty data is mileage, remaining fuel, longitude, or latitude, at least one set of co-clustered data of the first original data is determined from the multiple sets of third original data of the first vehicle; and In response to the determination that the data dimension of the dirty data is temperature, driver age, or driver gender, at least one set of co-clustered data of the first raw data is determined from the plurality of sets of second raw data of at least one second vehicle.

9. The processing method according to any one of claims 1 to 8, wherein, After performing the step of optimizing the first original data based on the at least one set of clustered data, the processing method further includes the following steps: Based on the optimized first raw data, corresponding recommendation information is provided to the user of the first vehicle.

10. A vehicle data processing device, characterized in that, include: Memory; as well as A processor, connected to the memory, and configured to implement the method for processing vehicle data as described in any one of claims 1 to 9.

11. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, the method for processing vehicle data as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Congestion area identification method based on traffic coil detection data quality control

    CN109754604A

  • User information processing method and device, computer equipment and storage medium

    CN110570229A

  • Data cooperative transmission method and device in VANET and electronic equipment

    CN113207080A