Automatic data management method and device and electronic equipment
By identifying abnormal situations in cloud data, and using TSP ciphertext data and various abnormal identification algorithms to automatically locate problem sources, the problem of low efficiency and accuracy of big data governance in the car cloud is solved, and the data quality problem positioning and user prompts are realized on the full link.
Patent Information
- Application Number
- CN202510745875.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the governance efficiency and accuracy of Cheyun big data is low, making it difficult to realize automated analysis and positioning of the full-link quality problems of data uploaded from the vehicle end to the cloud, making it difficult for data quality to meet actual needs.
By obtaining abnormal situations in cloud data, identifying the same period data in TSP ciphertext data, using multiple abnormal identification algorithms to locate the problem source, and sending prompt information to users, automatically locate the cause of the abnormal situation.
It improves the efficiency and accuracy of data governance, reduces the workload and error rate of users' manual troubleshooting of problem sources, and realizes the location and prompts of the full-link problem source of the data upload link.
Smart Images

Figure CN120263828A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicles, and particularly to a method, device and electronic device for automatic data governance. Background Art
[0002] With the development of automotive intelligence and networking, the amount of vehicle-cloud big data is growing explosively. The quality of vehicle-cloud big data is directly related to the realization of big data product functions, the effectiveness of vehicle remote control functions, the accuracy of battery big data monitoring, and the efficiency of after-sales problem troubleshooting. In the traditional vehicle-cloud big data governance mode, the identification and positioning of data quality problems highly rely on manual analysis. This method is not only inefficient but also difficult to guarantee accuracy. More prominently, in the process of data governance, it is difficult to conduct a complete quality problem analysis and positioning for the entire link from the vehicle end to the cloud end of the data, making it difficult for the data quality to meet the actual needs of vehicle-cloud big data.
[0003] Related technologies disclose a method for evaluating the average data quality of a single vehicle within a given statistical period by extracting single-vehicle sample single-day data and analyzing the outliers therein, so as to realize the data quality evaluation of the overall vehicle sample. Another related technology discloses an analysis of the data quality of electric vehicle charging equipment from multiple indicators such as accuracy, reliability, and timeliness, and establishes an overall evaluation method for electric vehicle data quality. It can be seen that most of the above methods are used to evaluate data quality. When there are quality problems, manual positioning is required, automation cannot be achieved, and the efficiency and accuracy of data governance are low. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device and electronic device for automatic data governance, aiming to solve the technical problem of low efficiency and accuracy of data governance.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows: In a first aspect, the embodiments of the present application provide a method for automatic data governance, the method including: obtaining cloud data uploaded by a vehicle to be inspected to a vehicle networking platform; in the case of identifying first data with an abnormal situation in the cloud data, obtaining second data in the ciphertext data of the in-vehicle intelligent service platform (Telematics Service Provider, TSP) of the vehicle to be inspected; wherein, the second data and the first data are in the same time period; by identifying whether there is an abnormal situation in the second data, positioning the problem source that causes the abnormal situation of the first data; based on the problem source, sending a prompt message to the user.
[0006] According to the above technical means, for the first data with abnormal conditions in the cloud data, obtain the second data in the same time period. By whether there are abnormal values in the second data, locate the problem source that causes the abnormal conditions, and send a prompt message to the user through the problem source. Considering the identification, location, and prompt of abnormal conditions in cloud data, automatically locate the problem source that causes abnormal conditions in the entire link from the data upload of the vehicle to be inspected to the cloud, more comprehensively and accurately locate the problem source and prompt the user, thereby improving the efficiency and accuracy of data governance.
[0007] In a possible implementation manner, by identifying whether there are abnormal conditions in the second data, locating the problem source that causes the abnormal conditions in the first data includes: when it is determined that there are no abnormal conditions in the second data, determining that the problem source is the data transmission path between the TSP of the vehicle to be inspected and the vehicle networking platform; when it is determined that there are abnormal conditions in the second data, determining that the problem source is the vehicle end.
[0008] According to the above technical means, by determining whether there are abnormal conditions in the second data, accurately locate the problem source of the abnormal conditions, thereby increasing the efficiency of data governance. In addition, through the automatic location of the problem source, reduce the workload and error rate of users' manual troubleshooting of the problem source, and improve the accuracy of data governance.
[0009] In a possible implementation manner, the data automatic governance method further includes: when it is determined that there are abnormal conditions in the second data, obtaining the third data in the vehicle end data of the vehicle to be inspected; wherein, the third data and the second data are in the same time period; when it is determined that there are no data quality problems in the third data, determining that the problem source is the data transmission path between the communication terminal of the vehicle to be inspected and the TSP; when it is determined that there are data quality problems in the third data, determining that the problem source is the controller in the vehicle to be inspected that generates the third data.
[0010] According to the above technical means, through the data quality problem acceptance of the third data, more accurately locate the problem source to the vehicle end controller or the data transmission path between the communication terminal of the vehicle and the TSP, further reduce the workload of users, and improve the efficiency and accuracy of data governance.
[0011] In a possible implementation manner, before obtaining the second data in the TSP data of the vehicle to be inspected when it is recognized that there is abnormal first data in the cloud data, the method further includes: using an anomaly recognition algorithm to identify whether there is data with abnormal conditions in the cloud data; wherein, the anomaly recognition algorithm includes at least one of the following: forward recognition algorithm, outlier recognition algorithm, anomaly value recognition algorithm, consistency acceptance algorithm, and integrity acceptance algorithm.
[0012] According to the above technical means, through a variety of anomaly recognition algorithms, data with abnormal conditions can be recognized more comprehensively and multi-dimensionally, significantly expanding the scope of application of the data automation governance method. The combination of multiple anomaly recognition algorithms can more accurately identify data with abnormal conditions, improving the accuracy of data automation governance.
[0013] In a possible implementation, the recognition rules of the forward recognition algorithm include at least one of the following categories: non-empty category, pattern matching category, status category, floating-point number category, and list category.
[0014] According to the above technical means, users can quickly configure the forward recognition algorithm for any field in the cloud data based on the five recognition rules of non-empty category, pattern matching category, status category, floating-point number category, and list category, so as to identify whether there is data with abnormal conditions in the cloud data, significantly reducing the complexity of data automation governance.
[0015] In a possible implementation, the outlier recognition rules of the outlier recognition algorithm include at least one of the following: single row and single column, single row and multiple columns, multiple rows and single column, multiple rows and multiple columns.
[0016] According to the above technical means, by configuring the outlier recognition algorithm based on the outlier recognition rules of single row and single column, single row and multiple columns, multiple rows and single column, and multiple rows and multiple columns, users can flexibly configure the outlier recognition rules, improving the flexibility of data automation governance.
[0017] In a possible implementation, when the anomaly recognition algorithm is an outlier recognition algorithm, the outlier recognition algorithm is used to identify whether there is data with abnormal conditions in the cloud data, including: grouping the cloud data to obtain at least one cloud data group; where the data in the same cloud data group is related; based on at least one cloud data group, constructing an isolation forest; for each cloud data in the cloud data group, traversing each isolation tree in the isolation forest, and obtaining the average height of the cloud data relative to the isolation forest based on the height of the cloud data relative to each isolation tree; based on the average height of the cloud data relative to the isolation forest, identifying whether the cloud data includes data with abnormal conditions.
[0018] According to the above technical means, grouping the cloud data through the correlation between data can significantly improve the efficiency and accuracy of identifying data with abnormal conditions. The outlier recognition algorithm based on the isolation forest can efficiently identify abnormal conditions in the cloud data through the construction of randomized isolation trees, with the characteristics of small resource consumption and high operation efficiency. The outlier recognition algorithm based on the isolation forest can process high-dimensional data and more accurately identify abnormal conditions, thus improving the efficiency and accuracy of the data automation governance method.
[0019] In a possible implementation, when the anomaly recognition algorithm is a consistency acceptance algorithm, the anomaly recognition algorithm is used to identify whether there is data with abnormal conditions in the cloud data, including: obtaining the vehicle-end data of the vehicle to be accepted; calculating the correlation coefficient between the vehicle-end data and the cloud data by using the consistency acceptance algorithm; and identifying whether there is data with abnormal conditions in the cloud data based on the correlation coefficient.
[0020] According to the above technical means, based on the correlation coefficient between the cloud data and the vehicle-end data, when there are abnormal conditions in the cloud data, users can be aware of it in time, avoiding the cloud providing incorrect services to the vehicle to be accepted based on the data with abnormal conditions, thus affecting the user's driving experience.
[0021] In a possible implementation, when the anomaly recognition algorithm includes an integrity acceptance algorithm, the anomaly recognition algorithm is used to identify whether there is data with abnormal conditions in the cloud data, including: dividing the cloud data into trips to determine a trip data group; obtaining a time difference set based on the time difference between two adjacent frames of cloud data in the trip data group; and identifying whether there is data with abnormal conditions in the cloud data based on the time difference set.
[0022] According to the above technical means, based on the time difference between two adjacent frames of cloud data, it is possible to identify data with jumps in the cloud data caused by reasons such as transmission or controllers, so that users can discover and handle abnormal conditions in time, significantly improving the efficiency of data automation governance.
[0023] In a possible implementation, dividing the cloud data into trips includes: dividing the cloud data based on the status of the vehicle to be accepted; or dividing the cloud data according to a preset time period.
[0024] According to the above technical means, by dividing the cloud data into trips in various ways such as the status of the vehicle to be accepted or a preset time period, the trip division method can be configured more flexibly, improving the flexibility of the data automation governance method.
[0025] In a second aspect, the present application provides a data automation governance device, including: an acquisition module, a processing module, and a sending module; the acquisition module is used to acquire the cloud data uploaded by the vehicle to be accepted to the vehicle networking platform; in the case of identifying the first data with abnormal conditions in the cloud data, acquiring the second data in the TSP ciphertext data of the vehicle to be accepted; where the second data and the first data are in the same time period; the processing module is used to locate the problem source that causes the abnormal conditions of the first data by identifying whether there are abnormal conditions in the second data; the sending module is used to send a prompt message to the user based on the problem source.
[0026] In a possible implementation, the processing module is specifically configured to determine that the problem source is the data transmission path between the TSP of the vehicle to be accepted and the vehicle networking platform when it is determined that the second data has no abnormal conditions; and determine that the problem source is the vehicle side when it is determined that the second data has abnormal conditions.
[0027] In a possible implementation, the processing module is further configured to obtain third data in the vehicle-side data of the vehicle to be accepted when it is determined that the second data has abnormal conditions; wherein, the third data and the second data are in the same time period; determine that the problem source is the data transmission path between the communication terminal of the vehicle to be accepted and the TSP when it is determined that the third data has no data quality problems; and determine that the problem source is the controller that generates the third data in the vehicle to be accepted when it is determined that the third data has data quality problems.
[0028] In a possible implementation, before obtaining the second data in the TSP data of the vehicle to be accepted when it is recognized that there is abnormal first data in the cloud data, the processing module is further configured to use an anomaly recognition algorithm to recognize whether there is data with abnormal conditions in the cloud data; wherein, the anomaly recognition algorithm includes at least one of the following: forward recognition algorithm, outlier recognition algorithm, anomaly value recognition algorithm, consistency acceptance algorithm, and integrity acceptance algorithm.
[0029] In a possible implementation, the recognition rules of the forward recognition algorithm include at least one of the following categories: non-empty category, pattern matching category, status category, floating-point number category, and list category.
[0030] In a possible implementation, the outlier recognition rules of the outlier recognition algorithm include at least one of the following: single row and single column, single row and multiple columns, multiple rows and single column, multiple rows and multiple columns.
[0031] In a possible implementation, when the anomaly recognition algorithm is the anomaly value recognition algorithm, the processing module is specifically configured to group the cloud data to obtain at least one cloud data group; wherein, the data in the same cloud data group is related; construct an isolation forest based on the at least one cloud data group; for each cloud data in the cloud data group, traverse each isolation tree in the isolation forest, and obtain the average height of the cloud data relative to the isolation forest based on the height of the cloud data relative to each isolation tree; and recognize whether there is data with abnormal conditions in the cloud data based on the average height of the cloud data relative to the isolation forest.
[0032] In a possible implementation, when the anomaly recognition algorithm is the consistency acceptance algorithm, the processing module is specifically configured to obtain the vehicle-end data of the vehicle to be accepted; calculate the correlation coefficient between the vehicle-end data and the cloud data by using the consistency acceptance algorithm; and identify whether there is abnormal data in the cloud data based on the correlation coefficient.
[0033] In a possible implementation, when the anomaly recognition algorithm includes the integrity acceptance algorithm, the processing module is specifically configured to divide the cloud data into trips to determine a trip data group; obtain a set of time differences based on the time differences between two adjacent frames of cloud data in the trip data group; and identify whether there is abnormal data in the cloud data based on the set of time differences.
[0034] In a possible implementation, the processing module is specifically configured to divide the cloud data based on the status of the vehicle to be accepted; or divide the cloud data according to a preset time period.
[0035] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the electronic device implements the method of the first aspect described above.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the data automation governance method provided in any of the embodiments of the first aspect described above is implemented.
[0037] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer program instructions, and when the computer program instructions are executed by a processor, the data automation governance method provided in any of the embodiments of the first aspect described above is implemented.
[0038] It should be noted that the technical effects brought by any implementation manner in the second aspect to the fifth aspect can refer to the technical effects brought by the corresponding implementation manner in the first aspect, and will not be elaborated here.
[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application and do not constitute an improper limitation to the present application.
[0041] Figure 1is a block diagram of a data automation governance system shown according to an exemplary embodiment; Figure 2 is a block diagram of a computing device shown according to an exemplary embodiment; Figure 3 is a flowchart of a data automation governance method shown according to an exemplary embodiment; Figure 4 is a flowchart of another data automation governance method shown according to an exemplary embodiment; Figure 5 is an architecture diagram of a data automation governance method shown according to an exemplary embodiment; Figure 6 is a flowchart of an outlier recognition algorithm shown according to an exemplary embodiment; Figure 7 is a block diagram of a data automation governance device shown according to an exemplary embodiment; Figure 8 is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0042] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0044] In the embodiments of the present application, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, article or device comprising the element.
[0045] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0046] For ease of understanding, the data automation governance method provided by the present application is specifically introduced below in conjunction with the accompanying drawings.
[0047] The data automation governance method provided by the present application can be applied to a data automation governance system as Figure 1 shown. The data automation governance system includes: a computing device 110, a vehicle to be inspected 120, a vehicle networking platform 130, and a TSP 140. Among them, the computing device 110 is respectively connected to the vehicle to be inspected 120, the vehicle networking platform 130, and the TSP 140.
[0048] In some embodiments, the vehicle to be inspected 120 and the vehicle networking platform 130 are communicatively connected through the TSP 140.
[0049] In some embodiments, as Figure 1 shown, the vehicle to be inspected 120 includes a controller 121 and a communication terminal 122.
[0050] As a feasible implementation manner, the controller 121 is used to generate vehicle-end data and send it to the communication terminal 122. The controller 121 can be a power system controller, a vehicle controller, a domain controller, etc. in the vehicle.
[0051] As another feasible implementation manner, the communication terminal 122 is used to receive the vehicle-end data and send it to the vehicle networking platform 130 after encryption through the TSP 140. Among them, the communication terminal 122 can be an in-vehicle terminal (Telematics BOX, T-BOX).
[0052] Exemplarily, the controller 121 and the communication terminal 122 are communicatively connected through a CAN bus (Controller Area Network).
[0053] As a feasible implementation manner, the vehicle networking platform 130 is a data storage, computing, and service center for the vehicle network.
[0054] In some embodiments, the computing device 110 is used to identify data with abnormal conditions, locate the problem source of the data that causes the abnormal conditions, and send a prompt message to the user. Exemplarily, the computing device 110 obtains the cloud data uploaded by the vehicle 120 to be inspected and accepted to the vehicle networking platform 130; in the case of identifying the first data with abnormal conditions in the cloud data, obtains the second data in the TSP ciphertext data of the vehicle 120 to be inspected and accepted; wherein, the second data and the first data are in the same time period; locates the problem source that causes the abnormal conditions of the first data by identifying whether there are abnormal conditions in the second data; and sends a prompt message to the user based on the problem source.
[0055] As a feasible implementation manner, the computing device 110 can be any device or equipment capable of performing data automated governance, such as a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or a computer. The embodiments of the present application do not make any limitations in this regard.
[0056] In some embodiments, as Figure 2 shown, the computing device 110 provided by the present application may include: a presentation layer, a service layer, and a data layer. Among them, the presentation layer is used to interact with the user, receive the positive recognition algorithm, the anomaly recognition algorithm, and the vehicle identification number (VIN) of the vehicle to be inspected and accepted input by the user, and send them to the service layer. The presentation layer can also send a prompt message to the user after the data quality problem recognition module determines that there is a data quality problem.
[0057] The service layer includes: an anomaly recognition algorithm input module, a data quality problem recognition module, and a data pulling module; the anomaly recognition algorithm input module is used to obtain the positive recognition algorithm and the anomaly recognition algorithm sent by the presentation layer, and send the positive recognition algorithm and the anomaly recognition algorithm to the data layer; the data quality problem recognition module is used to identify whether there are data quality problems in the cloud data based on the anomaly recognition algorithm; the data pulling module is used to determine the corresponding data based on the VIN number of the vehicle. Among them, the anomaly recognition algorithm includes: an outlier recognition algorithm, a consistency acceptance algorithm, an integrity acceptance algorithm, a positive recognition algorithm, and an anomaly recognition algorithm.
[0058] The data layer includes: an anomaly recognition algorithm library and a database; the anomaly recognition algorithm library is used to obtain the outlier recognition algorithm, the consistency acceptance algorithm, and the integrity acceptance algorithm, and receive the positive recognition algorithm and the anomaly recognition algorithm input by the anomaly recognition algorithm input module, and then send the anomaly recognition algorithm to the data quality problem recognition module; the database is used to obtain vehicle-side data, TSP ciphertext data, and cloud data, and send the corresponding data determined based on the data pulling module to the data quality problem recognition module.
[0059] It should be noted that the system architecture described in the embodiments of this application is for more clearly explaining the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of the system architecture, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0060] The data automation governance method provided by the embodiments of this application can be applied to Figure 1 the computing device 110 in the data automation governance system shown in the figure. As Figure 3 shown in the figure, the data automation governance method specifically includes the following steps: S301. Obtain the cloud data uploaded by the vehicle to be inspected to the vehicle networking platform.
[0061] As a feasible implementation method, based on the identification information of the vehicle to be inspected, obtain the corresponding cloud data in the vehicle networking platform. Among them, the identification information of the vehicle to be inspected can be the VIN number.
[0062] S302. In the case of identifying the first data with abnormal conditions in the cloud data, obtain the second data in the TSP ciphertext data of the vehicle to be inspected.
[0063] Among them, the second data and the first data are in the same time period.
[0064] As a feasible implementation method, during the process of the data being uploaded from the vehicle to be inspected to the vehicle networking platform, problems in any one of the controller of the vehicle to be inspected, the communication terminal of the vehicle to be inspected, the TSP, and the vehicle networking platform may cause the first data with abnormal conditions to appear in the cloud data in the vehicle networking platform. Therefore, in the case of identifying the first data with abnormal conditions in the cloud data, obtain the TSP ciphertext data at the same moment. The TSP ciphertext data, as the intermediate data between the vehicle-end data and the cloud data, records the original state of the vehicle-end data after leaving the vehicle to be inspected and can locate and trace the problem source that causes the abnormal conditions.
[0065] It should be understood that the TSP ciphertext data refers to the data used for storage or transmission after the TSP encrypts the vehicle-end data. The encryption process can significantly improve the security of the vehicle-end data during transmission and storage. It should be noted that when the TSP ciphertext data is sent to the vehicle networking platform, it will be temporarily stored in the Elasticsearch (ES). With the distributed architecture and real-time retrieval capabilities of the ES, the TSP ciphertext data can be quickly queried, analyzed, and visualized.
[0066] S303. By identifying whether there are abnormal conditions in the second data, locate the problem source that causes the abnormal conditions in the first data.
[0067] As a feasible implementation, in order to verify whether there are any abnormal situations when the vehicle - end data reaches the TSP, the anomaly recognition algorithm can be used to identify whether there are any abnormal situations in the second data.
[0068] It should be noted that since the second data is the data encrypted by the TSP, before performing the anomaly recognition, the second data needs to be decrypted.
[0069] S304. Send a prompt message to the user based on the problem source.
[0070] As a feasible implementation, through the positioning of the problem source, the corresponding engineer (i.e., the user) can be notified in a timely manner to handle the abnormal situation.
[0071] It can be understood that for the first data with abnormal situations in the cloud - end data, the second data in the same time period is obtained. By whether there are abnormal values in the second data, the problem source causing the abnormal situation is located, and a prompt message is sent to the user through the problem source. Considering the recognition, positioning, and prompting of abnormal situations in the cloud - end data, the problem source causing the abnormal situation is automatically located in the entire link from the data upload of the vehicle to be inspected to the cloud, more comprehensively and accurately locating the problem source and prompting the user, thereby improving the efficiency and accuracy of data governance.
[0072] In some embodiments, during the process of data uploading from the vehicle to be inspected to the vehicle networking platform, the problem source may include: the data transmission path between the vehicle end, the TSP, and the vehicle networking platform.
[0073] As a feasible implementation, when there are abnormal situations in the second data, it indicates that there have been abnormal situations when the data is uploaded from the vehicle to be inspected to the TSP, and the problem source causing the abnormal situation in the first data can be located as the vehicle end; when there are no abnormal situations in the second data, it means that there are abnormal situations when the data is uploaded from the TSP to the vehicle networking platform, and the problem source causing the abnormal situation in the first data can be located as the data transmission path between the TSP and the vehicle networking platform.
[0074] In some embodiments, when the problem source causing the abnormal situation in the first data is the vehicle end, the present application can further refine the problem source. For example, based on the generation and transmission path of the vehicle - end data, the vehicle end can be further refined to include: the data transmission path between the communication terminal of the vehicle to be inspected and the TSP, and the controller in the vehicle to be inspected. Therefore, the problem source causing the abnormal situation in the first data can be further determined from "the data transmission path between the communication terminal of the vehicle to be inspected and the TSP, and the controller in the vehicle to be inspected". Exemplarily, the specific positioning method can refer to the following steps S401 - S403, which will not be elaborated here.
[0075] In some embodiments, as Figure 4 shown, the above step S303 can be specifically implemented as the following steps: S3031. When it is determined that the second data has no abnormal situation, determine that the problem source is the data transmission path between the TSP of the vehicle to be inspected and the vehicle networking platform.
[0076] Exemplarily, if the second data has no abnormal situation, it indicates that the TSP has successfully received the vehicle-end data. However, if the first data has an abnormal situation, it indicates that an abnormality has occurred during the transmission of the data from the TSP to the vehicle networking platform. Therefore, it can be determined that the problem source is the data transmission path between the TSP of the vehicle to be inspected and the vehicle networking platform.
[0077] Exemplarily, when the problem source is located as the data transmission path between the TSP and the vehicle networking platform, the above step S304 can be specifically implemented as: sending an abnormal situation prompt message to the cloud big data platform development engineer.
[0078] S3032. When it is determined that the second data has an abnormal situation, determine that the problem source is the vehicle end.
[0079] Exemplarily, if the second data has an abnormal situation, it indicates that an abnormal situation has occurred when the data is uploaded from the vehicle to be inspected to the TSP. Then, the first data in the vehicle networking platform must have an abnormal situation. Therefore, it can be determined that the problem source is the vehicle end.
[0080] It can be understood that by determining whether the second data has an abnormal situation, the problem source of the abnormal situation is accurately located, thereby increasing the efficiency of data governance. In addition, through the automatic location of the problem source, the workload and error rate of users manually checking the problem source are reduced, and the accuracy of data governance is improved.
[0081] In some embodiments, when the problem source causing the abnormal situation of the first data is the vehicle end, the present application can further refine the problem source. For example, based on the generation and transmission path of the vehicle-end data, the vehicle end can be further refined to include: the data transmission path between the communication terminal of the vehicle to be inspected and the TSP, and the controller in the vehicle to be inspected. Therefore, the problem source causing the abnormal situation of the first data can be further determined from "the data transmission path between the communication terminal of the vehicle to be inspected and the TSP, and the controller in the vehicle to be inspected". Exemplarily, the specific location method can refer to the following steps S401-S403.
[0082] In some embodiments, to further locate the position of the problem source, the data automatic governance method provided by the present application further includes the following steps: S401. When it is determined that the second data is abnormal, obtain the third data in the vehicle-end data of the vehicle to be inspected and accepted.
[0083] Among them, the third data and the second data are in the same time period.
[0084] It should be understood that since the second data is encrypted after the vehicle-end data leaves the vehicle to be inspected and accepted, when it is recognized that the second data in the TSP is abnormal, in order to trace and locate the problem source that causes the abnormality of the second data, it is necessary to obtain the vehicle-end data at the same moment (i.e., the third data). By comparing the consistency between the third data and the second data, the problem source that causes the abnormal situation can be located.
[0085] Among them, the vehicle-end data can be vehicle-end CAN message data.
[0086] S402. When it is determined that the third data has no data quality problem, determine that the problem source is the data transmission path between the communication terminal of the vehicle to be inspected and accepted and the TSP.
[0087] As a feasible implementation method, it can be determined whether the third data has a data quality problem based on an anomaly recognition algorithm.
[0088] Exemplarily, when the third data has no data quality problem, it indicates that there is no problem with the controller in the vehicle to be inspected and accepted. However, since the second data is abnormal, it can be determined that there is a problem when the data is uploaded from the communication terminal of the vehicle to be inspected and accepted to the TSP, that is, it can be determined that the problem source is the data transmission path between the communication terminal of the vehicle to be inspected and accepted and the TSP.
[0089] Exemplarily, when the problem source is located as the data transmission path between the communication terminal of the vehicle to be inspected and accepted and the TSP, the above step S304 can be specifically implemented as: sending a prompt message to the T-box engineer.
[0090] S403. When it is determined that the third data has a data quality problem, determine that the problem source is the controller that generates the third data in the vehicle to be inspected and accepted.
[0091] Exemplarily, when the third data has a data quality problem, it indicates that the data has had a quality problem before being sent to the communication terminal. Therefore, it can be determined that the problem source is the controller that generates the third data in the vehicle to be inspected and accepted.
[0092] It should be noted that there are multiple controllers in the vehicle to be inspected. To further locate the controller with data quality problems, the third data can be mapped to each controller through the vehicle-side Matrix Protocol Matrix, and based on the real-time data corresponding to the third data in each controller, the controller with data quality problems can be accurately identified.
[0093] Exemplarily, in the case where the problem source is located at the controller that generates the third data in the vehicle to be inspected, the above S304 can be specifically implemented as: sending a prompt message to the software development engineer corresponding to the controller; so that the user can process it in time.
[0094] It can be understood that through the acceptance of the data quality problem of the third data, the problem source is more accurately located at the data transmission path between the vehicle-side controller or the vehicle's communication terminal and the TSP, further reducing the workload of the user and improving the efficiency and accuracy of data governance.
[0095] In some embodiments, before the above step S302, the data automation governance method provided by the present application further includes: using an anomaly recognition algorithm to identify whether there is data with abnormal conditions in the cloud data.
[0096] Among them, the anomaly recognition algorithm includes at least one of the following: forward recognition algorithm, discrimination recognition algorithm, outlier recognition algorithm, consistency acceptance algorithm, and integrity acceptance algorithm.
[0097] It should be noted that the forward recognition algorithm and the discrimination recognition algorithm can be configured by the user based on the actual situation of the vehicle.
[0098] It can be understood that through multiple anomaly recognition algorithms, data with abnormal conditions can be identified more comprehensively and multi-dimensionally, significantly expanding the applicable scope of the data automation governance method. The combination of multiple anomaly recognition algorithms can more accurately identify data with abnormal conditions and improve the accuracy of data automation governance.
[0099] In some embodiments, the forward recognition algorithm refers to determining whether the data is abnormal based on a preset recognition rule.
[0100] As a feasible implementation method, the recognition rules of the forward recognition algorithm include at least one of the following categories: non-empty category, pattern matching category, status category, floating-point number category, list category.
[0101] The non-empty category is used to represent that the cloud data is not a null value. Exemplarily, if the cloud data configured as the non-empty category is a null value at a certain moment, it is determined that the cloud data is data with abnormal conditions.
[0102] The pattern matching class is used to determine whether there are abnormal situations in the cloud data based on a preset field format. For example, the time data in the cloud data configured as the pattern matching class should satisfy the preset pattern matching regular expression r"^\d{4}-\d{2}-\d{2}\d{2}:\d{2}:\d{2}.*$", which stipulates that the time should be in the format of "year-month-day - hour:minute:second". If the time data in the cloud data configured as the pattern matching class fails to pass the matching verification of this regular expression, it indicates that there is an abnormality in the cloud data.
[0103] The status class is used to determine whether there are abnormal situations in the cloud data based on preset status values. Taking the door signal data as an example, if the preset status values of the door signal data are defined as "open" and "closed", when the door status shown in the cloud data configured as the status class does not belong to "open" or "closed", it is determined that there is an abnormality in the cloud data.
[0104] The floating-point number class is used to determine whether there are abnormal situations in the cloud data based on a preset range and preset precision requirements. Taking the battery voltage data of the vehicle to be inspected as an example, if the preset range is 0 - 1000V and the preset precision requirement is to retain 2 decimal places, then when the cloud data (i.e., the battery voltage data) configured as the floating-point number class exceeds this range (such as greater than 1000V) or does not meet the precision requirement (such as only retaining 1 decimal place), it is determined that the cloud data is abnormal data.
[0105] The list class is used to determine whether there are abnormal situations in the cloud data based on preset conditions. Taking the list of GPS track points of a vehicle to be inspected within a day as an example, the cloud data configured as the list class is a list containing longitude and latitude. The preset conditions include: the length of the list needs to be within a preset range (such as [100, 1000]), and each longitude and latitude coordinate needs to meet the geographical coordinate system specification. If the actual list data length of the cloud data configured as the list class is 99, it is determined that the list data is abnormal data.
[0106] It can be understood that users can quickly configure the positive recognition algorithm for any field in the cloud data based on these five recognition rules: non-empty class, pattern matching class, status class, floating-point number class, and list class, so as to identify whether there is abnormal data in the cloud data, significantly reducing the complexity of data automation governance.
[0107] In some embodiments, the outlier recognition rules of the outlier recognition algorithm include landmark data and enterprise standard data. After the cloud data triggers the outlier recognition rules, the cloud data at the corresponding moment will be recorded and statistically analyzed. The outlier recognition rules for landmark data are consistent with the regional supervision platform, which is used to ensure that the landmark data meets the regional requirements. The outlier recognition rules for enterprise standard data are set by users, which are used to ensure the stable operation of the vehicle networking platform.
[0108] As another feasible implementation method, the outlier recognition rules of the outlier recognition algorithm include at least one of the following: single row and single column, single row and multiple columns, multiple rows and single column, multiple rows and multiple columns.
[0109] Single row and single column is used to determine whether there is an abnormal situation in the cloud data based on a certain moment of a single field in the cloud data.
[0110] Single row and multiple columns is used to identify whether there is an abnormal situation in multiple cloud data based on the relevance of multiple cloud data at the same moment, where the multiple cloud data are relevant. Taking the multiple cloud data as vehicle speed data and vehicle status data as an example, when the vehicle speed data is greater than 0, if the vehicle status data is parked, it is determined that there is an abnormal situation in the multiple cloud data.
[0111] Multiple rows and single column is used to determine whether there is an abnormal situation in the cloud data based on the change trend of the cloud data. Taking the cloud data as the cumulative mileage of the vehicle to be inspected as an example, if the change trend of the cloud data is not monotonically increasing, it is determined that there is an abnormal situation in the cloud data.
[0112] Multiple rows and multiple columns is used to determine whether there is an abnormal situation in the cloud data based on the change trends of multiple cloud data. Taking multiple cloud data including accelerator pedal opening and vehicle speed as an example, when the accelerator pedal opening increases and the vehicle speed does not increase, it is determined that there is an abnormal situation in the cloud data.
[0113] It can be understood that by configuring the outlier recognition algorithm based on the outlier recognition rules of single row and single column, single row and multiple columns, multiple rows and single column, and multiple rows and multiple columns, users can flexibly configure the outlier recognition rules, which improves the flexibility of data automated governance.
[0114] In some embodiments, when the outlier recognition algorithm is an outlier value recognition algorithm, the step of "using the outlier recognition algorithm to identify whether there is abnormal data in the cloud data" in the above steps can be specifically implemented as the following steps: Sa1. Group the cloud data to obtain at least one cloud data group.
[0115] Among them, the data in the same cloud data group are relevant.
[0116] As a feasible implementation method, the relevance can be determined based on the physical meaning of the cloud data. The user groups the cloud data based on the physical meaning of the cloud data. Exemplarily, the cumulative mileage and the state of charge (SOC) of the battery are relevant and can be divided into the same cloud data group.
[0117] As another feasible implementation, the relevance can be determined based on the controller in the vehicle to be inspected of the cloud data. Exemplarily, the cloud data from the same controller is divided into a cloud data group.
[0118] It should be noted that the relevance can be determined based on any of the above methods or a combination of the two, and the embodiments of the present application do not limit this.
[0119] Sa2. Construct an isolation forest based on at least one cloud data group.
[0120] As a feasible implementation, assume that a certain cloud data group is: ; where is used to represent the cloud data group, is used to represent the nth cloud data, and n is used to represent the number of cloud data in the cloud data group.
[0121] Determine that the data of the ith cloud data in the cloud data group at any moment satisfies: ; The manifestation form is: , where is used to represent the ith cloud data, is used to represent the data of the ith cloud data at the dth moment.
[0122] Randomly select the cloud data at m moments from the cloud data at the above d moments to form a subset of X , and put into the root node of the isolation tree (i.e., the starting point of the isolation tree). It should be noted that randomly selecting the cloud data at m moments can reduce extreme complexity, and random selection reduces the risk of fitting and improves the efficiency of the outlier recognition algorithm.
[0123] Randomly specify a moment q from the cloud data at the above m moments, obtain multiple values of the cloud data at the qth moment, and randomly generate a cut point p between the maximum value and the minimum value among the multiple values of the cloud data, satisfying the following formula (1):
[0124] where is used to represent the cloud data at the qth moment, is used to represent the cloud data at m moments, is used to represent the cut point.
[0125] Exemplarily, a hyperplane is generated through the cutting point q, which can divide the cloud data at each moment into two subspaces and store them in the left and right nodes of the isolation tree respectively. For example, the cloud data less than p is put into the left node, and the cloud data greater than or equal to q is put into the right node. It should be noted that the cutting point is repeatedly selected to divide the cloud data until there is only one cloud data on all the leaf nodes of the isolation tree, or the isolation tree reaches the preset height, that is, the isolation tree is constructed.
[0126] As a feasible method, an isolation tree is constructed based on the cloud data group at each moment, and an isolation forest is composed of the isolation trees.
[0127] Sa3. For each cloud data in the cloud data group, traverse each isolation tree in the isolation forest, and based on the height of the cloud data relative to each isolation tree, obtain the average height of the cloud data relative to the isolation forest.
[0128] As a feasible implementation method, the height of the cloud data relative to each isolation tree is used to represent the path length from the root node to the cloud data at a certain moment. The average height is used to represent the average value of the height of each cloud data relative to the isolation tree at each moment.
[0129] Sa4. Based on the average height of the cloud data relative to the isolation forest, identify whether the cloud data includes data with abnormal conditions.
[0130] In some embodiments, based on the average height of the cloud data relative to the isolation forest, an anomaly index is determined; based on the anomaly index, identify whether there is data with abnormal conditions in the cloud data.
[0131] It should be understood that the anomaly index is used to measure the anomaly degree of abnormal data relative to normal data in the cloud data group. The higher the anomaly index, the more likely the cloud data has abnormal conditions.
[0132] Exemplarily, the anomaly index can satisfy the following formula (2):
[0133] Among them, is used to represent the anomaly index, is used to represent the average height, is used to represent the auxiliary function for calculating the anomaly index.
[0134] As a feasible implementation method, the anomaly index usually varies within the range of [0, 1). When the anomaly index is close to 1, it indicates a high probability of data with anomalies in the cloud data. When the anomaly index is close to 0, it indicates that there is no data with anomalies in the cloud data. For example, if the anomaly indexes of the cloud data in the same cloud data group are all close to 0.5, which is at the boundary of the abnormal situation, there will be no obvious abnormal situation.
[0135] Exemplarily, a preset anomaly threshold is set. When the anomaly index is greater than or equal to the anomaly threshold and the anomaly index of the cloud data is not near 0.5, the above cloud data is data with anomalies.
[0136] As another feasible implementation method, the auxiliary function of the anomaly index satisfies the following formula (3):
[0137] where, is used to represent the harmonic number from 1 to m - 1, and m is used to represent the number of moments of the cloud data. is used to represent the auxiliary function for calculating the anomaly index.
[0138] Exemplarily, the harmonic number from 1 to m - 1 can satisfy the following formula (4):
[0139] where, is used to represent the harmonic number from 1 to m - 1. is used to represent the natural logarithm function of m, and m is used to represent the number of moments of the cloud data.
[0140] It should be noted that the data with anomalies in each cloud data group are determined in sequence, and the number of data with anomalies and the corresponding timestamps in each cloud data group are counted to facilitate the user to handle the abnormal situation.
[0141] It can be understood that grouping the cloud data based on the correlation between data can significantly improve the efficiency and accuracy of identifying data with abnormal situations. The outlier identification algorithm based on Isolation Forest can efficiently identify abnormal situations in cloud data through the construction of randomized isolation trees. It has the characteristics of small resource consumption and high operation efficiency. The outlier identification algorithm based on Isolation Forest can process high-dimensional data and more accurately identify abnormal situations, thus improving the efficiency and accuracy of the data automated governance method.
[0142] In some embodiments, when the anomaly identification algorithm is the consistency acceptance algorithm, the step of "using the anomaly identification algorithm to identify whether there is data with abnormal situations in the cloud data" in the above steps can also be implemented as the following steps: Sc1. Obtain the vehicle-end data of the vehicle to be inspected and accepted.
[0143] As a feasible implementation method, obtain the vehicle-end data and cloud data of the vehicle to be inspected and accepted at the same time. The vehicle-end data is determined by parsing the vehicle-end CAN message data based on the Matrix protocol.
[0144] Sc2. Calculate the correlation coefficient between the vehicle-end data and the cloud data using a consistency acceptance algorithm.
[0145] As a feasible experimental method, the consistency acceptance algorithm includes: performing a time-series deduplication operation on the vehicle-end data and the cloud data to determine the vehicle-end deduplicated data group and the cloud-end deduplicated data group; determining the correlation coefficient based on the vehicle-end deduplicated data group and the cloud-end deduplicated data group.
[0146] Taking the vehicle-end data as an example, the time-series deduplication operation is used to represent removing duplicates from consecutive identical data in the vehicle-end data, only retaining one data, and at the same time removing the data in the vehicle-end data that has not been uploaded to the vehicle networking platform.
[0147] It should be understood that the correlation coefficient can be the Pearson correlation coefficient between the vehicle-end deduplicated data group and the cloud-end deduplicated data group. The Pearson correlation coefficient is a statistic used to measure the strength and direction of the linear relationship between two variables.
[0148] Exemplarily, assume that the vehicle-end deduplicated data group is , and the cloud-end deduplicated data group is ; where a is used to represent the number of vehicle-end deduplicated data and cloud-end deduplicated data, is used to represent the vehicle-end deduplicated data group, is used to represent the cloud-end deduplicated data, is used to represent the a-th vehicle-end deduplicated data, is used to represent the a-th vehicle-end deduplicated data.
[0149] In some embodiments, the correlation coefficient can satisfy the following formula (5):
[0150] Where, is used to represent the k-th vehicle-end deduplicated data, is used to represent the k-th cloud-end deduplicated data, is used to represent the mean value of the vehicle-end deduplicated data in the vehicle-end deduplicated data group, represents the mean value of the cloud-end deduplicated data in the cloud-end deduplicated data group.
[0151] Sc3. Based on the correlation coefficient, identify whether there is abnormal data in the cloud data.
[0152] As a feasible implementation, when preset conditions are met, determine the data in the cloud data that has abnormal conditions.
[0153] Among them, the preset conditions include at least one of the following: the correlation coefficient is less than the preset correlation threshold; the vehicle-side deduplicated data group does not contain all the cloud deduplicated data in the cloud deduplicated data group.
[0154] It should be understood that the preset correlation threshold is the lowest correlation coefficient for the vehicle-side deduplicated data and the cloud deduplicated data to maintain consistency. For example, the preset correlation coefficient can be 0.8.
[0155] It can be understood that based on the correlation coefficient between the cloud data and the vehicle-side data, when there are abnormal conditions in the cloud data, users can detect it in time, avoiding the cloud providing incorrect services to the vehicle to be inspected based on the data with abnormal conditions, thus affecting the user's driving experience.
[0156] In some embodiments, when the anomaly recognition algorithm includes an integrity acceptance algorithm, the step of "using the anomaly recognition algorithm to identify whether there is data with abnormal conditions in the cloud data" in the above steps can also be implemented as the following steps: Sd1. Divide the cloud data by trip to determine trip data groups.
[0157] As a feasible implementation, divide the cloud data based on the status of the vehicle to be inspected.
[0158] Exemplarily, the status of the vehicle to be inspected can be the driving status (the driving status can be determined based on the vehicle speed). For example, when the driving status of the vehicle to be inspected is the stationary state (the vehicle speed is 0) and the time in the stationary state is greater than or equal to the preset time threshold, divide the cloud data. Specifically, based on the trip between two adjacent stationary states of the vehicle, divide the cloud data uploaded to the vehicle networking platform into a trip data group.
[0159] As another feasible implementation, divide the cloud data according to a preset time period.
[0160] Exemplarily, when divided according to a preset time period, the time periods between different trip data groups are the same, and the operation of trip division is faster.
[0161] It can be understood that dividing the cloud data in multiple ways such as the status of the vehicle to be inspected or a preset time period can configure the trip division method more flexibly, improving the flexibility of the data automated governance method.
[0162] Sd2. Based on the time difference between two adjacent frames of cloud data in the trip data group, form a time difference set.
[0163] As a feasible implementation method, the time difference set can satisfy the following formula (6):
[0164] Wherein, is used to represent the time difference set, represents the time difference between the nth frame of cloud data and the (n + 1)th frame of cloud data.
[0165] Sd3. Identify the data with abnormal conditions in the cloud data based on the time difference set.
[0166] As a feasible implementation method, based on the time difference set, determine the parameter indicators; when the parameter indicators meet the preset parameter conditions, determine the data with abnormal conditions in the cloud data, where the parameter indicators include at least one of the following: mean, mode, median, variance, time difference de-duplication set.
[0167] Exemplarily, the preset parameter conditions may include at least one of the following: the mean is greater than or equal to the target mean; the mode and the median are not equal to the preset frequency threshold.
[0168] Wherein, the target mean can be determined based on the historical time difference mean, and the embodiments of the present application do not limit this.
[0169] The preset frequency threshold is the maximum frequency at which some time differences in the time difference set can appear, and the embodiments of the present application do not limit this.
[0170] It can be understood that based on the time difference between adjacent two frames of cloud data, it is possible to identify the data with jumps in the cloud data caused by reasons such as transmission or controller, so as to enable users to discover and handle abnormal situations in time, and significantly improve the efficiency of data automation governance.
[0171] In some embodiments, as Figure 5 shown, the data automation governance method provided by the present application receives the VIN number of the vehicle to be inspected input by the user and the configured abnormal identification algorithms (including: forward identification algorithm and outlier identification algorithm), obtains the corresponding cloud data in the vehicle networking platform through the VIN number of the vehicle to be inspected input by the user, and identifies the first data with abnormal conditions in the cloud data through the abnormal identification algorithms (including: forward identification algorithm, outlier identification algorithm, abnormal value identification algorithm, consistency acceptance algorithm, and integrity acceptance algorithm), obtains the second data (TSP ciphertext data), and locates the problem source that causes the abnormal conditions of the first data by identifying whether there are abnormal conditions in the second data.
[0172] Further, when the problem source is determined to be the data transmission path between the TSP and the vehicle networking platform, an exception prompt message is sent to the development engineer of the cloud big data platform; when the problem source is determined to be the data transmission path between the communication terminal of the vehicle to be accepted and the TSP, a prompt message is sent to the T-box engineer; when the problem source is determined to be the controller in the vehicle to be accepted, a prompt message is sent to the corresponding software development engineer of the controller, so as to realize the automated governance of data.
[0173] In some embodiments, as Figure 6 shown, the outlier recognition algorithm provided by this application can be implemented through the following steps: S601. Group the cloud data to obtain at least one cloud data group.
[0174] Among them, the data in the same cloud data group is relevant.
[0175] S602. Determine the root node.
[0176] Exemplarily, randomly select a cloud data group at a certain moment as the root node of the isolation tree.
[0177] S603. Select a cutting point.
[0178] Exemplarily, randomly select a cutting point from the cloud data group at a certain moment.
[0179] S604. Based on the cutting point, divide the cloud data at a certain moment into a left node and a right node.
[0180] S605. Continuously select cutting points to divide the cloud data group to construct an isolation tree.
[0181] S606. Construct an isolation forest based on the isolation tree.
[0182] S607. Determine the average height of the cloud data relative to the isolation forest.
[0183] S608. Based on the average height of the cloud data relative to the isolation forest, identify whether the cloud data includes data with abnormal conditions.
[0184] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, the data automation governance device or electronic device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0185] According to the above method, the embodiments of the present application can exemplarily divide the functional modules of the data automation governance device or electronic device. For example, the data automation governance device or electronic device can include each functional module corresponding to each functional division, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical functional division, and there can be other division methods in actual implementation.
[0186] Referring to Figure 7 , the data automation governance device 700 includes: an acquisition module 701, a processing module 702, and a sending module 703; the acquisition module 701 is used to acquire the cloud data uploaded by the vehicle to be inspected to the vehicle networking platform; in the case of identifying the first data with abnormal conditions in the cloud data, acquire the second data in the TSP ciphertext data of the vehicle to be inspected; wherein, the second data and the first data are in the same time period; the processing module 702 is used to locate the problem source that causes the abnormal conditions of the first data by identifying whether there are abnormal conditions in the second data; the sending module 703 is used to send a prompt message to the user based on the problem source.
[0187] In a possible implementation manner, the processing module 702 is specifically used to determine that the problem source is the data transmission path between the TSP of the vehicle to be inspected and the vehicle networking platform in the case of determining that the second data has no abnormal conditions; in the case of determining that the second data has abnormal conditions, determine that the problem source is the vehicle end.
[0188] In a possible implementation, the processing module 702 is further configured to, when it is determined that there is an abnormal situation in the second data, obtain third data in the vehicle-end data of the vehicle to be inspected; wherein, the third data and the second data are in the same time period; when it is determined that there is no data quality problem in the third data, determine that the problem source is the data transmission path between the communication terminal of the vehicle to be inspected and the TSP; when it is determined that there is a data quality problem in the third data, determine that the problem source is the controller that generates the third data in the vehicle to be inspected.
[0189] In a possible implementation, before obtaining the second data in the TSP data of the vehicle to be inspected when it is recognized that there is abnormal first data in the cloud data, the processing module 702 is further configured to use an anomaly recognition algorithm to recognize whether there is data with an abnormal situation in the cloud data; wherein, the anomaly recognition algorithm includes at least one of the following: forward recognition algorithm, outlier recognition algorithm, anomaly value recognition algorithm, consistency acceptance algorithm, and integrity acceptance algorithm.
[0190] In a possible implementation, the recognition rules of the forward recognition algorithm include at least one of the following categories: non-empty category, pattern matching category, status category, floating-point number category, list category.
[0191] In a possible implementation, the outlier recognition rules of the outlier recognition algorithm include at least one of the following: single row and single column, single row and multiple columns, multiple rows and single column, multiple rows and multiple columns.
[0192] In a possible implementation, when the anomaly recognition algorithm is the anomaly value recognition algorithm, the processing module 702 is specifically configured to group the cloud data to obtain at least one cloud data group; wherein, the data in the same cloud data group are related to each other; based on the at least one cloud data group, construct an isolation forest; for each cloud data in the cloud data group, traverse each isolation tree in the isolation forest, and based on the height of the cloud data relative to each isolation tree, obtain the average height of the cloud data relative to the isolation forest; based on the average height of the cloud data relative to the isolation forest, recognize whether there is data with an abnormal situation in the cloud data.
[0193] In a possible implementation, when the anomaly recognition algorithm is the consistency acceptance algorithm, the processing module 702 is specifically configured to obtain the vehicle-end data of the vehicle to be inspected; calculate the correlation coefficient between the vehicle-end data and the cloud data by using the consistency acceptance algorithm; based on the correlation coefficient, recognize whether there is data with an abnormal situation in the cloud data.
[0194] In a possible implementation, when the anomaly recognition algorithm includes an integrity acceptance algorithm, the processing module 702 is specifically configured to divide the cloud data into trips to determine a trip data group; obtain a time difference set based on the time difference between two adjacent frames of cloud data in the trip data group; and identify whether there is data with an anomaly in the cloud data based on the time difference set.
[0195] In a possible implementation, the processing module 702 is specifically configured to divide the cloud data based on the status of the vehicle to be accepted; or divide the cloud data according to a preset time period.
[0196] As Figure 8 shown, the electronic device 800 includes, but is not limited to, a processor 801 and a memory 802.
[0197] Among them, the above-mentioned memory 802 is used to store the executable instructions of the above-mentioned processor 801. It can be understood that the above-mentioned processor 801 is configured to execute instructions to implement the data automation governance method in the above-mentioned embodiments.
[0198] It should be noted that those skilled in the art can understand that Figure 8 the structure of the electronic device 800 shown in Figure 8 does not constitute a limitation on the electronic device. The electronic device 800 may include more or fewer components than
[0199] shown, or combine certain components, or have different component arrangements.
[0200] The processor 801 is the control center of the electronic device 800, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 802, and calling the data stored in the memory 802, it executes various functions of the electronic device 800 and processes data, thereby monitoring the electronic device 800 as a whole. The processor 801 may include one or more processing units. Optionally, the processor 801 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 801 either.
[0200] The memory 802 can be used to store software programs and various data. The memory 802 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one functional module (such as a determination unit, a processing unit, etc.). In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0201] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 802 including instructions. The above instructions can be executed by a processor 801 of an electronic device 800 to implement the data automation governance method in the above embodiment.
[0202] In actual implementation, Figure 7 the functions of the acquisition module 701, the processing module 702, and the sending module 703 in Figure 8 can all be implemented by the processor 801 calling a computer program stored in the memory 802. The specific execution process can refer to the description of the method part in the above embodiment and will not be elaborated here.
[0203] Optionally, the computer-readable storage medium can be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0204] In an exemplary embodiment, the embodiment of the present application also provides a computer program product including one or more instructions. The one or more instructions can be executed by a processor 801 of an electronic device 800 to complete the data automation governance method in the above embodiment.
[0205] It should be noted that when the instructions in the above computer-readable storage medium or the one or more instructions in the computer program product are executed by the processor of the electronic device 800, each process of the above method embodiment is implemented, and the same technical effects as the above method can be achieved. To avoid repetition, it will not be elaborated here.
[0206] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0207] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0208] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0209] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks or optical discs that can store program codes.
[0211] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for automated data governance, characterized in that The method includes: Obtaining cloud data uploaded by the vehicle to be inspected to the vehicle networking platform; When it is recognized that there is first data with an abnormal situation in the cloud data, obtaining second data in the TSP ciphertext data of the vehicle to be inspected; wherein, the second data and the first data are in the same time period; By identifying whether there is an abnormal situation in the second data, locating the problem source that causes the abnormal situation in the first data; Based on the problem source, sending a prompt message to the user.
2. The data automation governance method according to claim 1, wherein The step of locating the problem source that causes the abnormal situation in the first data by identifying whether there is an abnormal situation in the second data includes: When it is determined that there is no abnormal situation in the second data, determining that the problem source is the data transmission path between the TSP of the vehicle to be inspected and the vehicle networking platform; When it is determined that there is an abnormal situation in the second data, determining that the problem source is the vehicle end.
3. The data automation governance method according to claim 2, wherein The method further includes: When it is determined that there is an abnormal situation in the second data, obtaining third data in the vehicle end data of the vehicle to be inspected; wherein, the third data and the second data are in the same time period; When it is determined that there is no data quality problem in the third data, determining that the problem source is the data transmission path between the communication terminal of the vehicle to be inspected and the TSP; When it is determined that there is a data quality problem in the third data, determining that the problem source is the controller that generates the third data in the vehicle to be inspected.
4. The data automation governance method according to claim 1, wherein Before obtaining the second data in the TSP ciphertext data of the vehicle to be inspected when it is recognized that there is first data with an abnormal situation in the cloud data, the method further includes: Using an anomaly recognition algorithm to identify whether there is data with an abnormal situation in the cloud data; wherein, the anomaly recognition algorithm includes at least one of the following: forward recognition algorithm, discrimination recognition algorithm, outlier recognition algorithm, consistency acceptance algorithm, and integrity acceptance algorithm.
5. The data automation governance method according to claim 4, wherein The recognition rules of the forward recognition algorithm include at least one of the following categories: non-empty category, pattern matching category, status category, floating point number category, list category.
6. The data automation governance method according to claim 4, wherein The discrimination rules of the discrimination recognition algorithm include at least one of the following: single row and single column, single row and multiple columns, multiple rows and single column, multiple rows and multiple columns.
7. The data automation governance method according to claim 4, wherein When the anomaly recognition algorithm is an outlier recognition algorithm, the step of using the anomaly recognition algorithm to identify whether there is data with an abnormal situation in the cloud data includes: Grouping the cloud data to obtain at least one cloud data group; wherein, the data in the same cloud data group are related to each other; Based on the at least one cloud data group, constructing an isolation forest; For each cloud data in the cloud data group, traversing each isolation tree in the isolation forest, and obtaining the average height of the cloud data relative to the isolation forest based on the height of the cloud data relative to each isolation tree; Based on the average height of the cloud data relative to the isolation forest, identifying whether there is data with an abnormal situation in the cloud data.
8. The data automation governance method according to claim 4, wherein When the anomaly recognition algorithm is the consistency acceptance algorithm, the step of using the anomaly recognition algorithm to identify whether there is abnormal data in the cloud data includes: Obtain the in-vehicle data of the vehicle to be accepted; Calculate the correlation coefficient between the in-vehicle data and the cloud data using the consistency acceptance algorithm; Based on the correlation coefficient, identify whether there is abnormal data in the cloud data.
9. The data automation governance method according to claim 4, characterized in that When the anomaly recognition algorithm includes the integrity acceptance algorithm, the step of using the anomaly recognition algorithm to identify whether there is abnormal data in the cloud data includes: Perform trip division on the cloud data to determine trip data groups; Based on the time difference between two adjacent frames of cloud data in the trip data group, obtain a set of time differences; Based on the set of time differences, identify whether there is abnormal data in the cloud data.
10. The data automation governance method according to claim 9, wherein Performing trip division on the cloud data includes: Performing trip division on the cloud data based on the status of the vehicle to be accepted; or Performing trip division on the cloud data according to a preset time period.
11. A data automated governance device, characterized in that, It includes: An acquisition module, a processing module, and a sending module; The acquisition module is used to obtain the cloud data uploaded by the vehicle to be accepted to the vehicle networking platform; When it is recognized that there is first data with abnormal conditions in the cloud data, obtain second data in the TSP ciphertext data of the vehicle to be accepted; wherein, the second data and the first data are in the same time period; The processing module is used to locate the problem source that causes the abnormal situation of the first data by identifying whether there is an abnormal situation in the second data; The sending module is used to send a prompt message to the user based on the problem source.
12. An electronic device, characterized in that, It includes a processor and a memory, and the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the computer device to implement the data automation governance method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Data link monitoring method and device
CN105636100A
Channel fault point positioning method
CN111342893A
Tracing and positioning method and device for abnormal event of intelligent driving function of vehicle
CN114205223A
Abnormal node positioning method and device, electronic equipment and readable storage medium
CN115412430A
Internet of vehicles data quality monitoring method and system
CN117156473A