Data analysis method and device, equipment and storage medium
By weighted calculations of attribute credibility and usage credibility on shared data sets, the problem of determining data credibility in autonomous driving is solved, ensuring the authenticity and reliability of data, and promoting the healthy and efficient development of the data sharing environment.
Patent Information
- Application Number
- CN202510382302.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-01
AI Technical Summary
In autonomous driving technology, how to determine the credibility of data is an urgent problem to be solved, especially ensuring the authenticity of data during data sharing.
By obtaining the shared data set, analyzing multiple indicator values of the data object, weighted calculations based on the indicator value and predetermined weights, the attribute credibility and usage credibility of the data set are obtained, and weighted calculations are performed to obtain the overall credibility of the data set.
Ensure the authenticity of data, improve the health and efficiency of the data sharing environment, and inspire data providers to provide high-quality data.
Smart Images

Figure CN120408134A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data analysis method, apparatus, device, and storage medium. Background Art
[0002] With the rapid development of intelligence, the Internet of Vehicles (IoV) is becoming the focus of the automotive industry. Autonomous driving is the core of the IoV, and the development of autonomous driving technology depends on a large amount of real vehicle test data, including environmental perception data, decision-making control data, etc. during vehicle driving.
[0003] In related technologies of autonomous driving, whether it is the training of an autonomous driving model or the use of an autonomous driving model, a large amount of data is required, and a large amount of capital and effort are needed for data collection, annotation, and processing. In a scenario like autonomous driving that requires a large amount of diverse data, sharing data resources is inevitable.
[0004] However, in autonomous driving technology, safety is of utmost importance, and how to determine the credibility of data is an urgent problem to be solved. Summary of the Invention
[0005] To solve the above technical problems, this application provides a data analysis method, apparatus, device, and storage medium, and provides a method for analyzing data credibility to ensure that the data is true and reliable.
[0006] In a first aspect, this application provides a data analysis method, which includes: obtaining a shared data set, where the shared data set includes multiple data objects; analyzing the multiple data objects included in the shared data set to obtain the index values corresponding to each index of the shared data set; performing weighted calculation based on the index values corresponding to each index and the weights corresponding to each index determined in advance to obtain the attribute credibility of the shared data set; obtaining the usage information of the shared data set, and calculating the usage credibility of the shared data set based on the usage information; performing weighted calculation on the attribute credibility of the shared data set and the usage credibility of the shared data set to obtain the analysis result of the shared data set.
[0007] In a second aspect, the present application provides a data analysis device, which includes: a data set acquisition module, used to acquire a shared data set, which includes multiple data objects; an indicator value determination module, used to analyze the multiple data objects included in the shared data set to obtain indicator values corresponding to each indicator of the shared data set; an attribute credibility determination module, used to perform weighted calculation based on the indicator value corresponding to each indicator and the predetermined weight corresponding to each indicator to obtain the attribute credibility of the shared data set; a usage credibility determination module, used to acquire usage information of the shared data set and calculate the usage credibility of the shared data set based on the usage information; an analysis result determination module, used to perform weighted calculation on the attribute credibility of the shared data set and the usage credibility of the shared data set to obtain the analysis result of the shared data set.
[0008] In a third aspect, the present application provides a data analysis device, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by one or more processors, the one or more processors implement the data analysis method as described in the first aspect above.
[0009] In a fourth aspect, the present application provides a storage medium, which may be a computer-readable storage medium having a computer program stored thereon, and which, when executed by a processor, implements the data analysis method as described in the first aspect above.
[0010] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, implements any data analysis method as described in the first aspect above.
[0011] The technical solution provided by the embodiments of the present application has the following advantages compared with the prior art:
[0012] The embodiments of the present application provide a data analysis method, apparatus, device, and storage medium. After obtaining a shared data set, the method analyzes multiple data objects included in the shared data set to obtain indicator values corresponding to various indicators of the shared data set. A weighted calculation is performed based on the indicator values corresponding to the various indicators and predetermined weights corresponding to the various indicators to obtain the attribute credibility of the shared data set. Usage information of the shared data set is then obtained, and the usage credibility of the shared data set is calculated based on the usage information. Finally, a weighted calculation is performed on the attribute credibility of the shared data set and the usage credibility of the shared data set to obtain an analysis result of the shared data set. By considering both the attribute credibility and usage credibility of the data set, the overall credibility of the data set is obtained to ensure the authenticity and reliability of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0014] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1 It is a schematic structural diagram of a data analysis platform system provided by an embodiment of the present application;
[0016] Figure 2 It is a schematic flowchart of a data analysis method provided by an embodiment of the present application;
[0017] Figure 3 It is a schematic flowchart of calculating the attribute credibility of a shared data set provided by an embodiment of the present application;
[0018] Figure 4 A schematic diagram of the importance degree between any two dimensions provided by an embodiment of the present application;
[0019] Figure 5 A flowchart of a closed-loop system in the field of autonomous driving provided by an embodiment of the present application;
[0020] Figure 6 A structural diagram of the interaction between a data platform and a data closed-loop service system provided by an embodiment of the present application;
[0021] Figure 7 It is a schematic structural diagram of a data analysis device provided by an embodiment of the present application;
[0022] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0023] In order to more clearly understand the above objects, features, and advantages of the present application, the following will further describe the solutions of the present application. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0024] Many specific details are set forth in the following description to facilitate a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all of the embodiments.
[0025] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0026] It should be noted that the concepts such as "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0027] It should be noted that the modifications of "one" and "multiple" mentioned in this application are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0028] First, the application scenario of this application and the technical terms involved are introduced.
[0029] Autonomous driving data closed-loop: The autonomous driving data closed-loop refers to the process of continuously collecting, processing, analyzing and utilizing a large amount of data generated during the autonomous driving process to continuously optimize the algorithms and performance of the autonomous driving system. It includes key links such as data collection, annotation, simulation testing, model training, test verification and deployment. The data closed-loop is crucial for the continuous progress of the autonomous driving system because it allows the system to continuously learn through the data in actual driving to cope with more complex traffic conditions and long-tail problems.
[0030] Consortium blockchain: A cluster composed of multiple private blockchains, a blockchain jointly managed by multiple institutions. Each organization or institution manages one or more nodes, and its data is only allowed to be read, written and sent by different institutions within the system. The characteristics of the consortium blockchain include partial decentralization, strong controllability, data not being publicly available by default, and fast transaction speed. The consortium blockchain platform aims to promote cross-industry blockchain technology. The consortium blockchain is applicable to members of a specific group and limited third parties. Multiple preselected nodes are designated as bookkeepers within, and the generation of each block is jointly determined by all the preselected nodes. The maintenance and governance of the consortium blockchain are generally carried out by consortium members, usually adopting an election system, which is easy to perform permission control, and the code is generally partially open source or directionally open source.
[0031] Data element: It refers to the data resources that participate in social production and operation activities and bring economic benefits to the owners or users. The term "data element" is a reference to "data" in the context of discussing productivity and production relations in the digital economy, emphasizing the role of data in promoting production value. That is, data elements refer to computer data and its derivative forms that are aggregated, sorted, and processed according to specific production needs. The original data sets, standardized data sets, various data products, and systems, information, and knowledge generated based on data input into production can all be included in the scope of discussion of data elements.
[0032] "Corner case": A term commonly used in software engineering and system design, referring to special situations that occur under marginal or extreme conditions in system design. These conditions are usually not the most common ones but may occur in certain cases. If the system fails to handle these conditions properly, it may lead to errors or abnormal behaviors. In the field of autonomous driving, "Corner case" refers to those uncommon traffic scenarios that may pose challenges to autonomous driving systems.
[0033] In the field of autonomous driving, to ensure the security of shared data, blockchain technology has been introduced. The application of blockchain technology in autonomous driving data sharing features decentralization and immutability, which helps to define the rights and responsibilities of participants in each link.
[0034] According to the different degrees of centralization of blockchain networks, three different application scenarios are differentiated, namely public blockchains, consortium blockchains, and private blockchains.
[0035] A public blockchain refers to a blockchain that is completely open across the network without a user authorization mechanism. Anyone can join and participate in its activities. On a public blockchain, all transaction records are publicly transparent. Everyone can become a node of the network, and there is no central agency for control.
[0036] A consortium blockchain, also known as an industry blockchain, is a blockchain jointly managed by a group of organizations. Only pre-selected members have the right to join the network and participate in specific operations. Consortium blockchains are suitable for business scenarios where multiple entities need a shared and immutable ledger to record transactions but do not want any single party to control the entire system. Consortium blockchains offer higher privacy than public blockchains because transaction information is only shared among consortium members. At the same time, compared with private blockchains, consortium blockchains provide a certain degree of decentralization, reducing the risk of single points of failure.
[0037] A private blockchain is a completely closed blockchain owned and controlled by a single organization or individual. A private blockchain can be regarded as an internally used database, mainly used to improve efficiency, simplify processes, and enhance security. Although a private blockchain loses the decentralized feature of a public blockchain, it can better protect the privacy of sensitive data and allows for more efficient management and customized rule settings.
[0038] Since in the field of autonomous driving, multiple data providers are involved in sharing data resources, a consortium blockchain is used to realize the trustworthiness of shared data.
[0039] The process of data being uploaded to the blockchain usually involves three main stages: the pre-upload processing stage, the on-chain processing stage, and the smart contract processing stage. Each stage has its specific purpose and operations.
[0040] Specifically, in the pre-upload processing stage, the data to be uploaded needs to be preprocessed. This includes steps such as data collection, cleaning, verification, and formatting to ensure the quality and consistency of the data so that it can be accurately recorded on the blockchain.
[0041] The pre-upload processing stage also includes: the data encryption process to protect data privacy and security. Depending on the application scenario, different encryption algorithms may be used to ensure data security. Additionally, the pre-upload processing stage may also include the confirmation of data ownership or usage rights, and determining who has the right to upload this data to the blockchain.
[0042] The on-chain processing stage is mainly the process of actually writing the preprocessed data into the blockchain. This involves creating transactions (Transactions) and broadcasting these transactions to all nodes in the network. Once a transaction is broadcast, the nodes participating in the consensus mechanism will start to verify the validity of this transaction. After verification, the transactions will be packaged into blocks and finally added to the blockchain, completing the process of data being uploaded to the blockchain.
[0043] A smart contract is an automatically executed contract script directly written on the blockchain, which will be automatically triggered and executed when specific conditions are met. During the process of data being uploaded to the blockchain, if there are operations that need to be automatically executed according to certain rules, they can be implemented through smart contracts.
[0044] In the smart contract stage, the smart contract will further process the data uploaded to the blockchain. For example, it checks whether the data meets the requirements according to the pre-set logic, or automatically executes corresponding operations according to the data content. The execution result of the smart contract is also immutable, and the entire execution process is transparent, and all participants can view the content and execution status of the contract.
[0045] The above three stages together constitute the complete process of data uploading to the blockchain, ensuring that the entire process of data from preparation to final recording on the blockchain is secure, reliable, and traceable.
[0046] Meanwhile, using an architecture such as Figure 1 shown for the platform system enables the management of shared data.
[0047] As Figure 1 shown, it presents a data platform architecture for consortium blockchain management. The data platform is divided into three main parts: database and middleware, data management, and security management. The following is a detailed explanation of each part:
[0048] Among them, the database and middleware include: MySQL module, NoSQL module, and MQ module. Among them, the MySQL module is a relational database management system for storing structured data. The NoSQL module is a non-relational database suitable for handling large-scale data storage requirements, especially for semi-structured or unstructured data. The MQ module is a message queue (MessageQueue) for implementing asynchronous communication and decoupling system components.
[0049] Data management mainly includes: dataset management module, model management module, scenario library management module, alarm module, monitoring module, log module, data query module, event management module, data playback module, etc. Among them, the dataset management module is mainly responsible for the creation, maintenance, and management of datasets. The model management module is mainly used to manage data models to ensure data consistency and accuracy. The scenario library management module is mainly used to manage corresponding data sets according to different application scenarios. The alarm module is mainly used to monitor data status and trigger corresponding warning mechanisms. The monitoring module is mainly used to monitor the status and performance of the data platform in real time. The log module is mainly used to record operation logs for auditing and troubleshooting. The data query module is mainly used to provide data query interfaces to facilitate users to obtain the required data. The event management module is mainly used to manage events in the data platform, such as data change notifications. The data playback module is mainly used to support the historical playback function of data for data analysis and testing.
[0050] Security management mainly includes: operation audit module, security compliance module, key management module, and permission management module, etc. The operation audit module is mainly used to record all operation behaviors to ensure traceability. The security compliance module is mainly used to ensure that the data platform complies with relevant security standards and regulatory requirements. The key management module is mainly used to manage encryption keys to protect data security. The permission management module is mainly used to control the access permissions of different users to ensure secure access to data.
[0051] Figure 1The architecture diagram shown demonstrates how to manage and safeguard the data platform of the consortium blockchain through different modules to ensure its efficient and secure operation.
[0052] The data analysis process provided in the embodiments of this application is mainly carried out in the dataset management module of the above platform system architecture. The dataset management module includes operation functions such as data upload and data on-chain, as well as query functions such as data status and data analysis.
[0053] Next, in combination with the accompanying drawings and specific implementation manners, the data analysis method provided in the embodiments of this application will be described in detail.
[0054] Figure 2 It is a flowchart of a data analysis method in the embodiments of this application. This embodiment is applicable to the situation of analyzing shared datasets in the blockchain. This method can be executed by a data analysis device, which can be implemented in software and / or hardware, and the data analysis device can be configured in an electronic device.
[0055] As Figure 2 shown, the data analysis method provided in the embodiments of this application mainly includes steps S101 - S105.
[0056] S101. Obtain a shared dataset, where the shared dataset includes multiple data objects.
[0057] Among them, a dataset refers to a set of related data collected to meet specific analysis, modeling, or algorithm training requirements in the field of developing and testing autonomous driving technology. A shared dataset can be understood as a data set that has been uploaded to the consortium blockchain. The state of the shared dataset is the shared state or the on-chain state.
[0058] There are multiple data objects in the shared dataset, or it can be said that there are multiple data instances in the shared dataset. A data object refers to a specific record or sample representing a specific entity or object in a shared dataset. Each data object represents an independent unit in the dataset, and the data object can be any form of data, such as numerical values, text, images, audio, video, or point clouds, etc.
[0059] The shared dataset is used for training and simulation testing of the autonomous driving model.
[0060] In a possible implementation manner, the shared dataset includes at least one of the following types of data objects: video data, point cloud data, picture data, annotation information, sensor information.
[0061] Among them, the video data is a continuous image sequence of the road and the surrounding environment captured by a camera installed on a vehicle. The video data provides rich visual information for identifying pedestrians, other vehicles, traffic signs, etc. The attributes corresponding to the video data include at least one of the following: frame rate (FPS), resolution, timestamp, geographical location, shooting angle.
[0062] The point cloud data mainly comes from a Light Detection and Ranging (LiDAR), which generates a three-dimensional point cloud map of the surrounding environment by emitting laser beams and measuring the time of the reflected beams. The point cloud data is crucial for accurate distance measurement and obstacle detection. The attributes corresponding to the point cloud data include at least one of the following: point coordinates (X, Y, Z), intensity, number of echoes, timestamp, sensor position.
[0063] The picture data is similar to the video data, but refers to a single static image, which can be used for the analysis of a specific scene or moment. The picture data is particularly useful for detailed analysis and high-resolution imaging. The attributes corresponding to the picture data include at least one of the following: resolution, color depth, timestamp, shooting device model, geographical location.
[0064] The annotation information is a manual or automatic marking of the objects of interest in the original data (such as video, picture, point cloud). For example, the positions and categories of pedestrians and cars are annotated in the image. The annotation information is particularly important for training machine learning models. The attributes corresponding to the annotation information include at least one of the following: object category (such as pedestrian, car), bounding box coordinates, confidence score, labeler ID, timestamp.
[0065] The sensor information covers the data collected from various vehicle-mounted sensors (such as GPS, IMU, wheel speed sensors, etc.), providing information about the vehicle state and movement. The attribute values corresponding to the sensor information include at least one of the following: speed, acceleration, angular velocity, direction, geographical location (longitude, latitude), altitude, timestamp.
[0066] Obtain a shared dataset, including obtaining any dataset in any node on the consortium chain as the shared dataset.
[0067] S102. Analyze multiple data objects included in the shared dataset to obtain the metric values corresponding to each metric of the shared dataset.
[0068] The metrics of the shared dataset can be understood as the metrics used to evaluate the credibility of the basic information of the shared dataset preset in advance. The metric value refers to the specific value corresponding to the metric.
[0069] In a possible implementation, the metrics of the shared dataset include at least one of the following: clarity of image data, annotation accuracy of image data, authenticity of image data, scene label accuracy, accuracy of key event labels, scene integrity, key event integrity, data attribute integrity, multi-sensor time synchronization consistency, multi-sensor target feature consistency, data timeliness, geographical region diversity, weather condition diversity, traffic scene diversity, and so on.
[0070] All data in the shared dataset can be analyzed to obtain the metric values corresponding to the various metrics of the shared dataset. Alternatively, a part of the data can be randomly sampled for analysis to obtain the metric values corresponding to the various metrics of the shared dataset. This application embodiment does not specifically limit this.
[0071] Analyze the image data in the shared dataset to obtain the metric value corresponding to the clarity of the image data. Specifically, for at least one image data included in the shared dataset, use a no-reference image quality assessment algorithm to evaluate quality metrics such as the clarity of the image by extracting natural scene statistical features of the image, etc., and a quality score can be given without referring to the original clear image. Further, set a clarity metric threshold. If the image quality score is higher than this clarity metric threshold, it is considered to have high clarity. Finally, the proportion of the image data with an image quality score higher than this clarity metric threshold in the sampled image data volume is used as the metric value corresponding to the clarity of the image data. The metric value corresponding to the clarity of the image data is a value between 0 and 1.
[0072] Perform object detection and semantic segmentation annotation on the image data through an algorithm, and then compare the automatically annotated results with the annotation results provided by the user to calculate the intersection over union (IoU, Intersection over Union, IoU = intersection area / union area). Set an annotation accuracy threshold. For example, set the annotation accuracy threshold to 0.75. If the annotation accuracy is higher than this annotation accuracy threshold, it is considered that the image annotation is accurate. Finally, the proportion of the number of images with an annotation accuracy higher than this annotation accuracy threshold in the total number of annotated targets in the sample is used as the metric value corresponding to the annotation accuracy of the image data. The metric value corresponding to the annotation accuracy of the image data is a decimal between 0 and 1.
[0073] The metric value corresponding to the authenticity of the image data is determined through the following process: Determine whether the video and pictures are stitched or generated by generalization. Count the proportion of the number of images in the sample that are not determined to be stitched or generated by generalization in the sampled image number. The metric value corresponding to the authenticity of the image data is a decimal between 0 and 1.
[0074] The metric value corresponding to the accuracy of the scene label is determined through the following process: The model is used to identify the scene presented in the image and determine whether it is consistent with the provided scene label. For each scene, the proportion of the amount of data with accurate scene labels in the amount of sampled image data for that scene is counted; for all scenes in the dataset, the arithmetic mean of the accuracy rates for each scene is calculated as the final evaluation result. The metric value corresponding to the accuracy of the scene label is a decimal between 0 and 1.
[0075] The metric value corresponding to the accuracy of the key event label is determined through the following process: The evaluation criteria and process are similar to those of the scene label accuracy rate. The metric value corresponding to the accuracy of the scene label is a decimal between 0 and 1.
[0076] The metric value corresponding to the scene integrity is determined through the following process: The intelligent driving scenes are divided into major categories such as urban roads, highways, and rural roads, and each major category is further subdivided into different road conditions (such as congestion, smooth traffic, etc.) and weather conditions (such as sunny, rainy, snowy, etc.). The proportion of the scene categories covered by the data uploaded to the total number of predefined scene categories is counted. If the proportion is close to 100%, it indicates good integrity in terms of scene coverage and can provide sufficient samples for intelligent driving training. The metric value corresponding to the accuracy of the scene label is a decimal between 0 and 1.
[0077] The metric value corresponding to the key event integrity is determined through the following process: Similarly, a series of key intelligent driving events are predefined, such as emergency braking, lane change and overtaking, sudden appearance of pedestrians, etc. The proportion of the number of samples of key events to the number of samples of such events that should theoretically be covered is counted. The higher the proportion, the higher the data integrity, which helps the algorithm learn to handle various emergencies. The metric value corresponding to the accuracy of the scene label is a decimal between 0 and 1.
[0078] The metric value corresponding to the data attribute integrity is determined through the following process: Check whether it contains annotation data, necessary attribute information such as timestamps, geographical locations, vehicle states (speed, steering angle, etc.). The proportion of the number of data records that completely contain all specified attribute information to the total number of records is counted. The metric value corresponding to the accuracy of the scene label is a decimal between 0 and 1
[0079] The metric value corresponding to the multi-sensor time synchronization consistency is determined through the following process: Measure the matching degree of different sensors at the same moment. For example, calculate the distance deviation between the target position identified by the camera and the same target position detected by the radar at a certain time point, and count the mean and standard deviation of these deviations at each time point. Set a standard deviation threshold. If it is less than this threshold, it is considered that the multi-sensors are consistent in terms of time synchronization. The proportion of the number of data with consistent time synchronization to the total number of sensor data is counted. The metric value corresponding to the accuracy of the scene label is a decimal between 0 and 1.
[0080] The index values corresponding to the consistency of multi-sensor target features are determined through the following process: For the targets detected jointly, compare whether the target features obtained by different sensors (such as the appearance and color of the target obtained by the camera, the speed of the target obtained by the radar, etc.) are mutually corroborated and logically consistent. By setting the feature matching degree index (if the matching degree is above 90%, it is considered that the target features are consistent), count the proportion of the number of targets with feature matching in the total number of detected targets. As the final evaluation result. The index value corresponding to the accuracy of the scene label is a decimal between 0 and 1.
[0081] The index values corresponding to data timeliness are determined through the following process: Sort the collection times of all data in the dataset, take the median time as the collection time of the dataset, calculate the number of hours from this time to the current time, and then use the inverse proportional function as the mapping function, as shown in formula (1).
[0082]
[0083] Where c is a small positive number.
[0084] The mapping function can map the number of hours to a decimal between 0 and 1, and ensure that the larger the number of hours, the worse the timeliness and the lower the evaluation value.
[0085] The index values corresponding to geographical region diversity are determined through the following process: Count the number of different geographical regions covered by the data sources and the proportion of the data volume of each region in the total data volume. For example, if the data comes from multiple countries globally or multiple provinces and cities in the country, and the data distribution in each region is relatively uniform, it indicates that the data has good geographical region diversity.
[0086] The index values corresponding to weather condition diversity are determined through the following process: Determine the proportion of the data volume under different weather conditions (such as sunny, rainy, snowy, foggy, etc.). The closer the data proportion of each weather condition is to a uniform distribution, the better the diversity of the data in terms of weather.
[0087] The index values corresponding to traffic scene diversity are determined through the following process: Count the proportion of the data volume according to different traffic scene types (such as congested, unobstructed, construction sections, etc.).
[0088] The above diversity statistical data is quantified using Shannon entropy and the entropy value is normalized to the 0-1 interval.
[0089] S103. Perform weighted calculation based on the index values corresponding to each index and the weights corresponding to each index determined in advance to obtain the attribute credibility of the shared dataset.
[0090] Reliability refers to the reliability, authenticity, and trustworthiness of the shared dataset. Reliability reflects the level of trust in the accuracy and consistency of something. Attribute reliability can be understood as the reliability, authenticity, and trustworthiness of the shared dataset in terms of data attributes.
[0091] The weights corresponding to each indicator can be set according to the actual situation or given by experts with experience from all parties participating in the consortium chain. The embodiments of this application do not specifically limit it.
[0092] Perform weighted calculations on the indicator values corresponding to each indicator and the weights corresponding to each indicator to obtain a value, which is used as the attribute reliability of the shared dataset.
[0093] S104. Obtain the usage information of the shared dataset and calculate the usage reliability of the shared dataset based on the usage information.
[0094] Among them, the usage information of the shared dataset includes: the total number of times the shared dataset is used and the number of users who use the shared dataset.
[0095] Usage reliability can be understood as the reliability, authenticity, and trustworthiness of the shared dataset in terms of user usage.
[0096] The total number of times the shared dataset is used refers to the cumulative number of times the shared dataset is accessed, downloaded, or used for activities such as training models among all user members in the entire consortium chain network. Specifically, the access times refer to the number of times consortium chain members query or view the shared dataset. This includes, but is not limited to, browsing the dataset metadata. The download times represent the number of times consortium chain members download the dataset from the blockchain. More precisely, it should be the number of operations to obtain a copy of the dataset. The number of times used for training models refers to the specific number of times the shared dataset is used to train machine learning models. Each time a team or system uses this dataset for a new model training process, it will be recorded as one use. Other usage times: In addition to the above situations, if the dataset is also used for testing, verification, or other analysis purposes, these activities will also be included in the total number of uses.
[0097] The total number of times the shared dataset is used can assist in understanding the popularity of the dataset and its actual contribution to the research and development of autonomous driving technology.
[0098] The number of users using the shared dataset refers to the total number of different users who access, download, or utilize the shared dataset for various activities (such as training models, testing, etc.) within a specific time period. Specifically, the number of users accessing the shared dataset refers to the number of independent users who have viewed the dataset information (such as metadata, description documents, etc.) but not necessarily downloaded the actual data file. The number of users downloading the dataset refers to the actual number of users who have obtained a copy of the dataset from the consortium blockchain. Each independent account or entity's each download operation will be counted as an independent user's download behavior. The number of users using the dataset for analysis or training models includes all users who conduct further work based on this dataset, such as the number of users using it to train autonomous driving algorithms or other related applications. The number of users for other purposes refers to those users who, in addition to the above situations, use the dataset for other purposes (such as education, research, etc.), and these users will also be included in the total number of users.
[0099] The statistics of the number of users using the shared dataset helps to evaluate the value and influence scope of the shared dataset and better understand the effectiveness of data resources.
[0100] In a possible implementation, the usage credibility of the shared dataset is directly proportional to the total number of times the shared dataset is used; the usage credibility of the shared dataset is inversely proportional to the number of users using the shared dataset.
[0101] Specifically, the usage credibility of the shared dataset being directly proportional to the total number of times used means that as the total number of times the shared dataset is used increases, its usage credibility will also increase accordingly. As the total number of times the shared dataset is used decreases, its usage credibility will also decrease accordingly.
[0102] Specifically, a large number of usage times usually means there are more opportunities to verify the quality, accuracy, and applicability of the dataset. Each use can be regarded as a test of the dataset, and a large number of positive feedback can enhance users' trust in the dataset. If a shared dataset is frequently used, it may indicate that it has been recognized by multiple users or organizations, thus improving the overall credibility. A high usage rate may also prompt the provider to continuously update and improve the dataset, further enhancing its quality.
[0103] The inverse proportional relationship between the usage credibility of a shared dataset and the number of users using the shared dataset means that when the number of different users using the same shared dataset increases, the usage credibility of the shared dataset will instead decrease. Specifically, as the user group expands, users with different backgrounds, purposes, and skill levels start using the dataset, which may introduce more diverse interpretation methods and application scenarios. If these users do not correctly understand and apply the dataset, it may lead to misunderstandings or misuses, affecting the overall reputation of the dataset. More user access means higher security risks, including an increased likelihood of data leakage or improper use, which will also affect the credibility of the dataset. The participation of a large number of users may make it difficult to meet the consistency requirements of the dataset, especially when there are differences in the data formats, contents, etc. expected by users.
[0104] In a possible implementation, calculating the usage credibility of a shared dataset based on usage information includes: taking the ratio of the total number of times the shared dataset is used to the number of users using the shared dataset as the initial usage credibility of the shared dataset; performing a normalization calculation on the initial usage credibility of the shared dataset to obtain the usage credibility of the shared dataset.
[0105] Calculate the initial usage credibility of the shared dataset through formula (2).
[0106]
[0107] The initial usage credibility reflects the usage frequency of the dataset by each user on average. A higher initial usage credibility may indicate that the shared dataset has higher practicality or value because users tend to use it multiple times; a lower initial usage credibility may imply that there may be certain limitations in the dataset or it does not fully meet the expectations of users.
[0108] If a dataset is frequently used by a small number of users (high total usage times but low number of users), it may mean that these users have a high dependence on the dataset, but it may also be because the dataset is highly specialized and only applicable to specific fields or scenarios.
[0109] On the contrary, if a large number of users only use the shared dataset a small number of times, it may indicate that although the dataset has attracted wide attention, it fails to fully meet the needs of most users, or users only make a preliminary attempt and do not use it deeply.
[0110] Furthermore, in order to more accurately reflect the comparability between shared datasets of different scales and consider the differences in different application scenarios, the initial usage credibility can be normalized.
[0111] Specifically, perform the normalization through formula (3).
[0112]
[0113] The purpose of normalization is to convert the initial usage credibility into a score between 0 and 1, where a value close to 1 represents higher credibility, and a value close to 0 represents lower credibility.
[0114] In the embodiments of the present application, not only the overall usage frequency of the dataset is considered, but also the user base of actual usage is taken into account, thus providing a more comprehensive perspective for evaluating the quality of the dataset.
[0115] S105. Perform weighted calculation on the attribute credibility of the shared dataset and the usage credibility of the shared dataset to obtain the analysis result of the shared dataset.
[0116] As described above, the analysis result of the shared dataset consists of two parts: the attribute credibility of the shared dataset and the usage credibility of the shared dataset, and the value ranges of both the attribute credibility and the usage credibility are [0, 1]. Among them, the attribute credibility of the shared dataset is triggered for calculation when the data is uploaded to the chain and does not change over time. While the usage credibility of the shared dataset changes over time.
[0117] And over time, the index proportion of the usage credibility of the shared dataset in the analysis result will be higher and higher. Therefore, the following formula (4) is defined to represent the analysis result of the shared dataset.
[0118] Analysis result of the dataset = α × attribute credibility + (1 - α) × usage credibility (4)
[0119] Among them, the value range of α is [0, A], and A can be set as needed, with the value range of (0, 1). The proportion of the usage credibility of the dataset increases over time and reaches the set maximum value A at a certain moment, and then the proportion no longer changes.
[0120] The embodiments of the present application provide a data analysis method. After obtaining the shared dataset, based on the fact that the shared dataset includes multiple data and at least one attribute value corresponding to each data in the multiple data, determine the index values corresponding to each index of the shared dataset; perform weighted calculation based on the index values corresponding to each index and the weights corresponding to each index determined in advance to obtain the attribute credibility of the shared dataset; then obtain the usage information of the shared dataset and calculate the usage credibility of the shared dataset based on the usage information; finally, perform weighted calculation on the attribute credibility of the shared dataset and the usage credibility of the shared dataset to obtain the analysis result of the shared dataset. Since starting from two aspects of the attribute credibility and the usage credibility of the dataset, the overall credibility of the dataset is obtained to ensure the authenticity and reliability of the data.
[0121] In a possible implementation, the data provider is incentivized based on the analysis results of the shared dataset.
[0122] Incentivizing the data provider based on the analysis results of the shared dataset means designing and implementing corresponding reward mechanisms according to the analysis results of aspects such as the attributes of the shared dataset and its usage effects, so as to encourage the data provider to continue to provide high-quality data or improve its data provision behavior.
[0123] This can not only enhance the enthusiasm of the data provider, but also promote a healthier and more efficient data sharing environment.
[0124] Based on the above embodiments, the embodiment of the present application further optimizes the calculation process of the attribute credibility of the shared dataset, as Figure 3 shown. The calculation method of the attribute credibility of the optimized shared dataset mainly includes steps S1031 - S1034.
[0125] S1031. Determine the dimension to which each index of the shared dataset belongs based on the pre-determined correspondence between the indexes and dimensions.
[0126] In the embodiment of the present application, multiple dimensions of the dataset are preset, including but not limited to: data accuracy, data integrity, data consistency, data timeliness, and data diversity.
[0127] Data accuracy refers to the degree to which the data correctly reflects the real-world situation, that is, the accuracy of the data with respect to the real object or event it describes. Data integrity means that the data is complete without any important information missing. Data consistency means that the data characteristics collected by multiple sensors are consistent. Data diversity means that the dataset covers statistical data of different scenarios, different weather conditions, and different regions.
[0128] The pre-determined correspondence between the indexes and dimensions is as follows: The data accuracy dimension includes the following indexes: clarity of image data, annotation accuracy of image data, authenticity of image data, scene label accuracy, and accuracy of key event labels. The data integrity dimension includes: scene integrity, key event integrity, and data attribute integrity. The data consistency dimension includes: multi-sensor time synchronization consistency and multi-sensor target feature consistency. The data timeliness dimension includes: data timeliness. The data diversity dimension includes: geographical region diversity, weather condition diversity, and traffic scene diversity.
[0129] Determine the dimensions corresponding to each indicator according to the pre-determined correspondence between indicators and dimensions. For example: the clarity of image data, the annotation accuracy of image data, the authenticity of image data, the scene label accuracy, and the accuracy of key event labels. The dimensions corresponding to these 5 indicators are the data accuracy dimension. The dimensions corresponding to these 3 indicators: scene integrity, key event integrity, and data attribute integrity are the data integrity dimension. The dimensions corresponding to these 2 indicators: multi-sensor time synchronization consistency and multi-sensor target feature consistency are the data integrity dimension. The dimension corresponding to the indicator of data timeliness is data timeliness. The dimensions corresponding to these 3 indicators: geographical region diversity, weather condition diversity, and traffic scene diversity are the data diversity dimension.
[0130] S1032. Determine the weights corresponding to each dimension.
[0131] Regarding the credibility of the attribute information of the data set, weighted averaging can be performed according to the weights of the dimensions. The weight of each dimension can be obtained by the priority calculation method of the Analytic Hierarchy Process (AHP).
[0132] Specifically, define the dimensions that affect the credibility of data attributes, including the above 5 parts, namely: data accuracy, data integrity, data consistency, data timeliness, and data diversity.
[0133] Determine the weights corresponding to each dimension, including: determining the priority parameters of each dimension; constructing a decision matrix based on the priority parameters of each dimension; constructing a matrix equation based on the n-order identity matrix and the decision matrix. The matrix equation is represented by formula (5):
[0134] |λE - A| = 0 (5)
[0135] Among them, E represents the n-order identity matrix, A represents the decision matrix, and λ represents the unknown parameter; solve the matrix equation to obtain the maximum eigenvalue of the unknown parameter; use the maximum eigenvalue as the weight corresponding to each dimension.
[0136] Make pairwise comparisons for all dimensions to compare the priorities between any two dimensions. The priority parameter represents the importance degree between 2 dimensions. As Figure 4 shown, the definitions of 1 - 9 are as follows: 1 represents equal importance, 3 represents medium importance, 5 represents importance, 7 represents very important, 9 represents extremely important, the importance of 2 is between 1 and 3, the importance of 4 is between 3 and 5, the importance of 6 is between 5 and 7, and the importance of 8 is between 7 and 9.
[0137] Figure 4The importance levels between any two dimensions are given. For example, data accuracy is very important relative to data integrity; data accuracy is important relative to data consistency; data accuracy is very important relative to data timeliness; data accuracy is extremely important relative to data diversity. Data integrity is equally important relative to data consistency; data integrity is moderately important relative to data timeliness; data integrity is between important and very important relative to data diversity. Data consistency is between equally important and moderately important relative to data timeliness; data consistency is moderately important relative to data diversity. Data timeliness is between equally important and moderately important relative to data diversity.
[0138] Figure 4 The values of the importance comparison between the two dimensions in ij can be given by experienced experts from all parties participating in the consortium chain, and the final priority parameters are obtained by taking the average.
[0139] Furthermore, a decision matrix is constructed based on the priority parameters of each dimension.
[0140] According to Figure 4 In the priority comparison in ij , a decision matrix is constructed. There are 5 dimensions: data accuracy, data integrity, data timeliness, data consistency, and data diversity. Then, a 5×5 decision matrix is constructed. Then, according to Figure 4 In the priority parameters in ij , it is filled into the decision matrix. Each element a ij in the decision matrix represents the importance ratio of dimension i relative to dimension j. The decision matrix is shown in Table 1.
[0141] Table 1
[0142] Data accuracy Data integrity Data consistency Data timeliness Data diversity Data accuracy 1 7 5 7 9 Data integrity 0.14 1 1 3 6 Data consistency 0.2 1 1 2 3 Data timeliness 0.14 0.33 0.5 1 2 Data diversity 0.11 0.17 0.33 0.5 1
[0143] It should be noted that when constructing the decision matrix, logical errors may occur, so it is necessary to use consistency checking to see if there are any problems.
[0144] In addition, the experts from all parties participating in the evaluation only intervene to give professional opinions when the standards change, and the specific scores thereafter should also be automatically updated and calculated by the system.
[0145] Subsequently, the weight values of each dimension are obtained according to the decision matrix, and the weight values are determined based on the principal eigenvector of the decision matrix.
[0146] Specifically, let the above decision matrix be A. The main eigenvalue λ is obtained by formula (5) as the modulus maximum value, and X obtained by substituting it into (λE - A)X = 0 is the main eigenvector. The main eigenvector is (60.9%, 15.6%, 12.7%, 6.7%, 4.1%), which is the weight corresponding to each dimension. The weight of the data accuracy dimension is 60.9%, the weight of the data integrity dimension is 15.6%, the weight of the data consistency dimension is 12.7%, the weight of the data timeliness dimension is 6.7%, and the weight of the data diversity dimension is 4.1%.
[0147] S1033. Determine the weights corresponding to the various indicators of the shared data set based on the dimensions to which the various indicators of the shared data set belong and the weights corresponding to each dimension.
[0148] Directly use the weights of each dimension as the weights corresponding to the indicators included in that dimension. The data accuracy dimension includes the following indicators: the clarity of image data, the annotation accuracy of image data, the authenticity of image data, the scene label accuracy, and the accuracy of key event labels. Then, the weights of these 5 indicators, namely, the clarity of image data, the annotation accuracy of image data, the authenticity of image data, the scene label accuracy, and the accuracy of key event labels, are 60.9%.
[0149] The data integrity dimension includes: scene integrity, key event integrity, and data attribute integrity. The weights of these 3 indicators, namely, scene integrity, key event integrity, and data attribute integrity, are 15.6%.
[0150] The data consistency dimension includes: multi-sensor time synchronization consistency, and multi-sensor target feature consistency. The weights of these 2 indicators, namely, multi-sensor time synchronization consistency and multi-sensor target feature consistency, are 12.7%.
[0151] The data timeliness dimension includes: data timeliness. The weight corresponding to this indicator of data timeliness is 6.7%.
[0152] The data diversity dimension includes: geographical region diversity, weather condition diversity, and traffic scene diversity. The weights of these 3 indicators, namely, geographical region diversity, weather condition diversity, and traffic scene diversity, are 4.1%.
[0153] S1034. Perform weighted calculation based on the indicator values corresponding to the various indicators and the weights corresponding to the various indicators to obtain the attribute credibility of the shared data set.
[0154] After obtaining the weights corresponding to each index in S1033, the index values corresponding to each index and the weights corresponding to each index are weighted and calculated, and the obtained value is used as the attribute credibility of the shared data set. Exemplarily, 0.845 calculated in Table 2 is the attribute credibility of the shared data set.
[0155] Table 2
[0156]
[0157] In the embodiment of the present application, the AHP priority algorithm is used to calculate the attribute credibility of the shared data set, which can not only effectively integrate multiple information sources, but also improve the quality of the attribute credibility.
[0158] It should be noted that in the embodiment of the present application, multiple indicators are divided into 5 dimensions, and the AHP priority algorithm is used to calculate the weights of each dimension to optimize the calculation process and reduce the calculation process.
[0159] In practical applications, the AHP priority algorithm can be used to calculate the weights corresponding to each indicator respectively to improve the calculation accuracy and further improve the quality of the credibility of the shared data set.
[0160] On the basis of the above embodiments, the embodiment of the present application provides a closed-loop system for data processing, management and model training in the field of autonomous driving. This closed-loop system not only covers the whole process from data collection to model deployment, but also particularly emphasizes the importance of data security, standardization and evaluation mechanism. Specifically, as Figure 5 shown, it mainly includes the following processes:
[0161] S201. Collect data on autonomous driving standards.
[0162] Collect data from vehicle terminals and roadside devices. Vehicle terminals include but are not limited to: sensors, cameras, etc.; roadside devices include but are not limited to: traffic lights, surveillance cameras, etc.
[0163] The goal is to obtain comprehensive and diverse driving environment information to provide a basis for subsequent analysis.
[0164] S202. Process the collected data.
[0165] Before uploading to the cloud, the data must be decrypted and desensitized to protect personal privacy and sensitive information.
[0166] Standardized format conversion ensures that all data has consistency and compatibility, facilitating subsequent processing and analysis.
[0167] S203. Upload the original data and chain the data.
[0168] The preprocessed data is securely stored in the blockchain, ensuring the immutability and transparency of the data.
[0169] S204. Annotate the original data.
[0170] The annotation process includes automatic annotation and manual annotation, aiming to provide accurate learning materials for the machine learning model. Automatic annotation includes: using algorithms to automatically identify and label objects. Manual annotation includes: manual inspection and correction.
[0171] S205. Upload the annotated data and put the annotated data on the chain.
[0172] The annotated data is also recorded in the blockchain to ensure its integrity and traceability.
[0173] S206. Use the annotated data for model training.
[0174] Use high-quality annotated data to train the autonomous driving algorithm to improve the accuracy and robustness of the model.
[0175] S207. Use simulation data to test and verify the trained model.
[0176] Use the simulation environment to test the performance of the model and possibly further verify the model performance by combining real-world data not used for training.
[0177] S208. According to the test situation, determine whether the model meets the deployment conditions. If so, execute S209; if not, execute S206.
[0178] Evaluate the performance of the model under various conditions and decide whether it can be deployed to actual vehicles for use.
[0179] S209. Conduct inference verification on the deployed model.
[0180] Run the model in a real environment, monitor its performance, and ensure that it can stably and reliably execute the expected tasks.
[0181] Combined with the above data closed-loop processing flow method, the overall business process is jointly completed by the data platform and the data closed-loop business system. The platforms are independent of each other and cooperate with each other. The system modules and business processes are as Figure 6 shown.
[0182] The data platform is responsible for core functions such as data collection, processing, storage, putting on the chain, and data analysis. The business system focuses on specific autonomous driving application requirements, such as data annotation, model training, simulation testing, etc.
[0183] The data analysis method evaluates the credibility of a data set through quantitative indicators (such as the ratio of the total number of uses to the number of users and its normalization), which is crucial for ensuring the safety and effectiveness of autonomous driving technology. This solution constructs a complete closed-loop system for data processing and model development, paying attention to both the management and security of the data itself and evaluating the data value through scientific methods, thus promoting the development of autonomous driving technology.
[0184] Figure 7 It is a schematic structural diagram of a data analysis device in an embodiment of the present application. As Figure 7 shown, the data analysis device 70 provided in the embodiment of the present application mainly includes: a data set acquisition module 71, configured to acquire a shared data set, where the shared data set includes multiple data objects; an index value determination module 72, configured to analyze the multiple data objects included in the shared data set to obtain the index values corresponding to the respective indexes of the shared data set; an attribute credibility determination module 73, configured to perform weighted calculation based on the index values corresponding to the respective indexes and the weights corresponding to the respective indexes determined in advance to obtain the attribute credibility of the shared data set; a usage credibility determination module 74, configured to acquire the usage information of the shared data set and calculate the usage credibility of the shared data set based on the usage information; an analysis result determination module 75, configured to perform weighted calculation on the attribute credibility of the shared data set and the usage credibility of the shared data set to obtain the analysis result of the shared data set.
[0185] In a possible implementation manner, the usage information of the shared data set includes: the total number of uses of the shared data set and the number of users using the shared data set; the usage credibility of the shared data set is in a direct proportional relationship with the total number of uses of the shared data set; the usage credibility of the shared data set is in an inverse proportional relationship with the number of users using the shared data set.
[0186] In a possible implementation manner, the usage credibility determination module 74 is specifically configured to use the ratio of the total number of uses of the shared data set to the number of users using the shared data set as the initial usage credibility of the shared data set; perform normalization calculation on the initial usage credibility of the shared data set to obtain the usage credibility of the shared data set.
[0187] In a possible implementation manner, the attribute credibility determination module 73 is specifically configured to determine the dimension to which each index of the shared data set belongs based on the correspondence between the indexes and dimensions determined in advance; determine the weights corresponding to each dimension; determine the weights corresponding to the respective indexes of the shared data set based on the dimension to which each index of the shared data set belongs and the weights corresponding to each dimension; perform weighted calculation based on the index values corresponding to the respective indexes and the weights corresponding to the respective indexes to obtain the attribute credibility of the shared data set.
[0188] In a possible implementation, the attribute credibility determination module 73 is specifically configured to obtain the priority parameters of each dimension; construct a decision matrix based on the priority parameters of each dimension; construct a matrix equation based on the n-order identity matrix and the decision matrix, and the matrix equation is represented by the following formula:
[0189] |λE - A| = 0
[0190] where E represents the n-order identity matrix, A represents the decision matrix, and λ represents the unknown parameter; solve the matrix equation to obtain the main eigenvector; and use the main eigenvector as the weight corresponding to each dimension.
[0191] In a possible implementation, the incentive module is configured to incentivize the data provider based on the analysis result of the shared data set.
[0192] In a possible implementation, the shared data set includes at least one of the following types of data objects: video data, point cloud data, picture data, annotation information, and sensor information; the shared data set is used to train and simulate test the autonomous driving model.
[0193] The data analysis device provided by the embodiments of the present application can execute the data analysis method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0194] Figure 8 It is a schematic structural diagram of an electronic device provided in this embodiment. The electronic device may include a data analysis device, as Figure 8 shown, the electronic device 800 includes a processor 810, a memory 820, an input device 830, and an output device 840; the number of processors 810 in the electronic device may be one or more, Figure 8 and one processor 810 is taken as an example here; the processor 810, the memory 820, the input device 830, and the output device 840 in the electronic device may be connected by a bus or other means, Figure 8 and the connection by bus is taken as an example here.
[0195] The memory 820, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data analysis method in the embodiments of the present invention. The processor 810 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 820, that is, implements the data analysis method provided by the embodiments of the present invention.
[0196] The memory 820 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 820 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 820 may further include a memory remotely provided with respect to the processor 810, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0197] The input device 830 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device, and may include a keyboard, a mouse, etc. The output device 840 may include a display device such as a display screen.
[0198] This embodiment also provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to implement the data analysis method provided by the embodiments of the present invention when executed by a computer processor.
[0199] Certainly, for the storage medium containing computer-executable instructions provided by the embodiments of the present invention, the computer-executable instructions are not limited to the above method operations, and can also execute relevant operations in the data analysis methods provided by any embodiments of the present invention.
[0200] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.
[0201] It should be noted that in the embodiments of the above data analysis device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0202] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0203] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data analysis method, characterized in that, The method includes: Obtaining a shared data set, which includes a plurality of data objects; Analyzing the plurality of data objects included in the shared data set to obtain the metric values corresponding to the respective metrics of the shared data set; Performing weighted calculation based on the metric values corresponding to the respective metrics and the weights corresponding to the respective metrics determined in advance to obtain the attribute credibility of the shared data set; Obtaining the usage information of the shared data set and calculating the usage credibility of the shared data set based on the usage information; Performing weighted calculation on the attribute credibility of the shared data set and the usage credibility of the shared data set to obtain the analysis result of the shared data set.
2. The method according to claim 1, characterized in that, The usage information of the shared data set includes: the total number of times the shared data set is used and the number of users using the shared data set; The usage credibility of the shared data set is in a direct proportional relationship with the total number of times the shared data set is used; The usage credibility of the shared data set is in an inverse proportional relationship with the number of users using the shared data set.
3. The method according to claim 2, wherein The calculating the usage credibility of the shared data set based on the usage information includes: Taking the ratio of the total number of times the shared data set is used to the number of users using the shared data set as the initial usage credibility of the shared data set; Performing normalization calculation on the initial usage credibility of the shared data set to obtain the usage credibility of the shared data set.
4. The method according to any one of claims 1 to 3, characterized in that The performing weighted calculation based on the metric values corresponding to the respective metrics and the weights corresponding to the respective metrics determined in advance to obtain the attribute credibility of the shared data set includes: Based on the pre-determined correspondence between metrics and dimensions, determining the dimensions to which the respective metrics of the shared data set belong; Determining the weights corresponding to the respective dimensions; Based on the dimensions to which the respective metrics of the shared data set belong and the weights corresponding to the respective dimensions, determining the weights corresponding to the respective metrics of the shared data set; Performing weighted calculation based on the metric values corresponding to the respective metrics and the weights corresponding to the respective metrics to obtain the attribute credibility of the shared data set.
5. The method according to claim 4, wherein The determining the weights corresponding to the respective dimensions includes: Obtaining the priority parameters of the respective dimensions; Constructing a decision matrix based on the priority parameters of the respective dimensions; Constructing a matrix equation based on the n-order identity matrix and the decision matrix, the matrix equation is represented by the following formula: |λE - A| = 0 where E represents the n-order identity matrix, A represents the decision matrix, and λ represents the unknown parameter; Solving the matrix equation to obtain the principal eigenvector; Taking the principal eigenvector as the weights corresponding to the respective dimensions.
6. The method according to claim 1, wherein It further includes: Motivating the data provider based on the analysis result of the shared data set.
7. The method according to claim 1, characterized in that The shared data set includes at least one of the following types of data objects: video data, point cloud data, picture data, annotation information, sensor information; the shared data set is used for training and simulation testing of an autonomous driving model.
8. A data analysis device, characterized in that, The apparatus includes: A data set acquisition module for obtaining a shared data set, which includes a plurality of data objects; An index value determination module, configured to analyze multiple data objects included in the shared data set to obtain index values corresponding to respective indexes of the shared data set; An attribute credibility determination module, configured to perform weighted calculation based on the index values corresponding to respective indexes and weights corresponding to respective indexes determined in advance, to obtain the attribute credibility of the shared data set; A usage credibility determination module, configured to obtain usage information of the shared data set and calculate the usage credibility of the shared data set based on the usage information; An analysis result determination module, configured to perform weighted calculation on the attribute credibility of the shared data set and the usage credibility of the shared data set to obtain an analysis result of the shared data set.
9. A data analysis device, characterized in that, The device includes: One or more processors; A storage device, configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data analysis method according to any one of claims 1-7.
10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the data analysis method according to any one of claims 1-7.