Data processing method and related equipment

By extracting and storing the data to be processed, the problems of low processing efficiency and high resource consumption in the prior art are solved, and data compression and processing efficiency are improved.

CN120067624APending Publication Date: 2025-05-30CONTEMPORARY AMPEREX FUTURE ENERGY RES INST (SHANGHAI) LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311607636.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, due to the large amount of original data and the diverse and complex data forms, the data reading and calculation cost for information evaluation is high, the processing efficiency is low, and the resource consumption is high.

Method used

By introducing data indicator types and data feature levels, feature extraction of data to be processed can be obtained, and data features at multiple levels are stored on the feature platform. This method reduces the access requirement for raw data, improves data processing efficiency, and reduces resource consumption.

Benefits of technology

It realizes a large amount of data to be processed, reduces the complexity of data utilization with large data volume and diverse and complex data forms, reduces the cost of data storage and data reading calculation, improves data processing efficiency, and reduces resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067624A_ABST
    Figure CN120067624A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of data processing, and provides a data processing method and related equipment, and the method comprises the steps: obtaining to-be-processed data and a data index type; according to the data index type, performing feature extraction on the to-be-processed data to obtain data features of different levels; and storing the data features to a feature platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a data processing method and related devices. Background Art

[0002] The core of a data product is the flow and consumption of data. This requires information evaluation based on a large amount of raw data, and then notifying or presenting the evaluation results to users. For example, in battery data management, after fault evaluation and health evaluation of the battery data uploaded from the vehicle end, the evaluation results are fed back to the users for notification or warning.

[0003] However, raw data usually has the characteristics of large data volume and diverse and complex data forms. The access and utilization of raw data result in high data reading and computing costs, low processing efficiency, and high resource consumption. Summary of the Invention

[0004] Embodiments of this application provide a data processing method and related devices to solve the problems in the prior art that due to the large data volume and diverse and complex data forms of raw data, the data reading and computing costs for information evaluation are high, the processing efficiency is low, and the resource consumption is high.

[0005] In a first aspect of embodiments of this application, a data processing method is provided, including:

[0006] Obtain data to be processed and data index types;

[0007] According to the data index types, perform feature extraction on the data to be processed to obtain data features at different levels;

[0008] Store the data features in a feature platform.

[0009] In the above data processing process, when performing feature extraction on the data to be processed, data index types and data feature levels are introduced. Multiple levels of data features are extracted based on the data index types, enabling the data to be processed to be represented in the form of multi-level data features in combination with the data index types, ensuring the diversity and multi-level nature of the features when representing the data to be processed in the form of data features, and being able to achieve a large-scale compression of the data to be processed, reducing the utilization complexity of the data to be processed with a large data volume and diverse and complex data forms, reducing the data storage volume and data reading and computing costs, helping to improve the data processing efficiency, and reducing resource consumption.

[0010] In some embodiments, the performing feature extraction on the data to be processed according to the data index types to obtain data features at different levels includes:

[0011] Extract features from the data to be processed according to the first data index type to obtain data features at the first level;

[0012] According to the dependency relationships among different data index types, sequentially perform feature extraction of the subsequent data index type on the data features at the current level obtained by extraction, to obtain data features at the subsequent level, until data features at multiple levels are obtained.

[0013] In the above process, by using the dependency relationships among data index types, based on this dependency relationship, sequentially perform feature processing on the data to be processed. Through layer-by-layer progressive feature processing, different-level data features are obtained. During the process of obtaining different-level data features, repeated access to the data to be processed is reduced, the feature processing efficiency is improved, the data reading and calculation cost is reduced, and resource consumption is decreased.

[0014] In some embodiments, performing feature extraction of the subsequent data index type on the data features at the current level obtained by extraction to obtain data features at the subsequent level includes:

[0015] Aggregate the data features at the current level according to the target data granularity to obtain aggregated data; the target data granularity is greater than the data granularity of the data features at the current level;

[0016] Perform feature extraction on the aggregated data according to the subsequent data index type to obtain data features at the subsequent level.

[0017] In this process, the already generated data features are used as the data basis for generating features of the subsequent data index type. Due to the granularity differences between the data index types with dependency relationships and different-level data features, the feature contents at different levels are aggregated, and feature extraction is performed on the aggregated data based on the subsequent data index type to obtain data features at the subsequent level, so as to realize the layer-by-layer progressive processing of feature contents in a more effective way and improve the feature processing efficiency.

[0018] In some embodiments, after storing the data features in the feature platform, it further includes:

[0019] Query the first required features adapted to the algorithm project from the feature platform;

[0020] When it is determined that the first required features exist in the feature platform, input the first required features into the algorithm project to obtain the feature processing result output by the algorithm project.

[0021] The above process changes the way of directly accessing the data to be processed, utilizes the different-level data feature extraction requirements obtained in advance and stored in the feature platform for algorithm engineering utilization, realizes the sharing and common use of data features through the feature platform, improves the feature reuse rate, reduces the data processing complexity of evaluation calculation, improves the data processing efficiency, and reduces resource consumption.

[0022] In some embodiments, after querying the first requirement feature adapted to algorithm engineering from the feature platform, it further includes:

[0023] In the case where it is determined that the first requirement feature does not exist in the feature platform, obtain the selected data as the data to be processed, and return to execute the step of extracting different-level data features from the data to be processed according to the data index type until it is determined that the first requirement feature exists in the feature platform.

[0024] In the above processing process, in the case where the requirement feature is queried from the feature platform but not found, it can be determined that the selected data is used as the data to be processed, so as to perform data feature processing on the selected data to generate different-level data features, ensure that the current requirement feature is stored in the feature platform, realize data expansion and system black start, and facilitate the effective development of subsequent information evaluation processing.

[0025] In some embodiments, after storing the data features in the feature platform, it further includes:

[0026] Select a second requirement feature from the feature platform and directly output the second requirement feature to the user side.

[0027] This process provides a feature output method, bypasses the information evaluation calculation processing of algorithm engineering, directly outputs the data features obtained by prior processing in the feature platform to the user, meets the feature output requirements of the data to be processed in specific situations, reduces the data evaluation processing flow, and expands the application scenario and information processing efficiency of the feature data in the feature platform.

[0028] In some embodiments, the method further includes:

[0029] Select target data content from the data to be processed and directly output the target data content to the user side.

[0030] This process provides a data output method, directly outputs the appropriate data in the data to be processed to the user, meets the data output requirements in specific situations, and expands the data application scenario and information processing efficiency.

[0031] In some embodiments, the method further includes:

[0032] Evaluate the stored data features in the feature platform;

[0033] In the case of determining that the stored data features need to be iterated, execute the step of extracting features from the data to be processed according to the data index type to obtain data features at different levels.

[0034] This process ensures the timely iterative update of the data stored in the feature platform and the effectiveness of the features through the evaluation of the stored data features in the feature platform.

[0035] In some embodiments, the method further includes:

[0036] Evaluate the stored data features in the feature platform;

[0037] In the case of determining that there are target data features to be taken offline in the stored data features, take offline the target data features from the feature platform.

[0038] This process ensures the timely offline update of the data stored in the feature platform and the effectiveness of the features through the evaluation of the stored data features in the feature platform.

[0039] A second aspect of the embodiments of the present application provides a data processing device, including:

[0040] A data acquisition module, configured to acquire data to be processed and a data index type;

[0041] A feature generation module, configured to extract features from the data to be processed according to the data index type to obtain data features at different levels;

[0042] A feature storage module, configured to store the data features in a feature platform.

[0043] A third aspect of the embodiments of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0044] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0045] The fifth aspect of the present application provides a computer program product, including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes the steps in the method described in the first aspect above.

[0046] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And in all the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0048] Figure 1 is a flowchart of the data processing method in some embodiments of the present application;

[0049] Figure 2 is an example diagram of the data feature matrix in some embodiments of the present application;

[0050] Figure 3 is a flowchart of the data processing method in some embodiments of the present application;

[0051] Figure 4 is an example diagram of a dependency relationship between data index types in some embodiments of the present application;

[0052] Figure 5 is a flowchart of the data processing method in some embodiments of the present application;

[0053] Figure 6 is a flowchart of the data processing method in some embodiments of the present application;

[0054] Figure 7 is a flowchart of the data processing method in some embodiments of the present application;

[0055] Figure 8 is an overall example diagram of the implementation process of the data processing method in some embodiments of the present application;

[0056] Figure 9 is a structural diagram of the data processing device in some embodiments of the present application;

[0057] Figure 10 is a structural diagram of the computer device provided in the embodiments of the present application. Detailed implementation manners

[0058] The embodiments of the technical solutions of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solutions of the present application more clearly, and therefore are only examples and cannot be used to limit the protection scope of the present application.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion.

[0060] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "a plurality" is more than two unless otherwise specifically defined.

[0061] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0062] In the description of the embodiments of this application, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0063] The flow and consumption of data are the core of data products.

[0064] In actual situations, taking the battery cloud management platform as an example, its core is to use the uploaded battery data as the original data, directly perform fault assessment and health assessment on these original data through algorithm engineering, and then notify or present the assessment results to the user through user engineering to implement battery information notification or warning.

[0065] With the increase in battery usage time, the expansion of battery data sources, and other factors, the continuous upload of battery data is huge. Each device uploads thousands of data items per day, resulting in a large amount of data obtained and complex and diverse data forms. When a large number of algorithms are used in the information evaluation process, it will also trigger a batch access to the original data, making the data reading and calculation costs for accessing and using the original data high, the processing efficiency low, and the resource consumption high.

[0066] In an embodiment of the present application, a data processing method is provided, enabling the data to be processed to be characterized in the form of multi-level data features, so as to reduce the access to the original data, reduce resource consumption, and improve data processing efficiency.

[0067] In order to illustrate the technical solutions described in the present application, specific embodiments will be used for illustration below.

[0068] Combined with Figure 1 As shown, in some embodiments, a data processing method is proposed, including:

[0069] Step 101, obtain the data to be processed and the data index type.

[0070] The data to be processed may specifically be original data, source data, etc. Or it may be existing data generated by other means, such as different-level data features that have already been generated.

[0071] Data indicators can be used to measure the quality of data, statistical trends, degree of change, etc. Data indicators have different types.

[0072] Data index types can be divided into the following several types:

[0073] Qualitative indicators: Indicators used to describe or represent specific features, attributes, or properties, usually non-numerical. For example, color, gender, brand, etc.

[0074] Quantitative indicators: Indicators used to represent quantities, quantity relationships, or numerical values. Usually measurable, numerical calculations and statistical analyses can be performed. For example, age, income, sales volume, etc.

[0075] Continuous indicators: Indicators with an infinite number of possible values, which can vary continuously within a range. For example, time, temperature, height, etc.

[0076] Discrete indicators: Indicators with a finite number of possible values, which cannot vary continuously within a range. For example, the number of students, the number of products, grades, etc.

[0077] Relative indicators: Indicators used to compare or relatively evaluate the differences or relationships between different entities. For example, market share, growth rate, ratio, etc.

[0078] Absolute indicators: Indicators that exist independently and do not depend on other indicators or reference points. For example, total population, total sales, etc.

[0079] Basic indicators: The original data that serves as the basis for other indicators and is used to calculate or derive other indicators. For example, total revenue, total cost, etc.

[0080] Derived indicators: Indicators calculated or derived from basic indicators. For example, profit margin, market growth rate, etc.

[0081] These data indicator types can be further subdivided and classified according to specific data analysis requirements and fields.

[0082] Step 102: Extract features from the data to be processed according to the data indicator type to obtain data features at different levels.

[0083] There is a set corresponding relationship between the data indicator type and different levels.

[0084] The relationship between the data indicator type and the data feature level can be a one-to-one relationship, or a many-to-one relationship, or a one-to-many relationship.

[0085] During the implementation process, multiple levels of data features can be obtained by extracting features from the data to be processed according to the same data indicator type; or multiple levels of data features can be obtained by extracting features from the data to be processed according to multiple data indicator types; or a certain level of data features can be obtained by extracting features from the data to be processed according to multiple data indicator types, and multiple levels of data features can be obtained in this way.

[0086] Data features can be hierarchically divided based on the feature relationships between features, so that different data features have corresponding levels. Among them, the feature relationships here are, for example, independent parallel relationships, or size relationships in terms of feature granularity, or feature dependency relationships.

[0087] Correspondingly, the different levels of data features can be independent parallel hierarchical relationships. For example, if one data indicator type corresponds to one feature dimension, then the data features at one level extracted under each data indicator type also correspond to one feature dimension, and the data feature levels form a parallel hierarchical relationship. As shown in Figure 2 shown, the corresponding data features can be presented in the form of a feature matrix.

[0088] Alternatively, the different levels of data features can have a non-parallel hierarchical relationship with forward and backward dependencies. For example, different levels respectively correspond to data features with different feature granularities on the same feature dimension. The feature granularity can be time granularities such as days, weeks, months, etc., or population granularities such as individuals, families, communities, etc. Then, corresponding data feature hierarchies will be generated among the data features with different feature granularities.

[0089] In the process of performing feature extraction on the data to be processed in step 102, it is necessary to introduce feature engineering to enable the extraction and processing of features from the data to be processed, so as to obtain feature data.

[0090] Feature engineering is the process of extracting features that can better represent the essential features of the problem from the data to be processed.

[0091] Feature engineering can perform data preprocessing and transformation on the basis of the data to be processed to extract more useful features, so that subsequent algorithms (such as machine learning algorithms) can better understand and utilize these features. The goal of feature engineering is to enable subsequent algorithms to better learn patterns and rules from the data by selecting, constructing, and transforming features. The main tasks of feature engineering include data cleaning, feature selection, feature construction, and feature transformation, etc.

[0092] With the help of feature engineering, data feature extraction is performed based on the data to be processed, generating data features at multiple levels, enabling the data to be processed to be represented in the form of multi-level data features, which can reduce the data volume by several times or even dozens of times, achieve a large-scale compression of the data to be processed, greatly reduce the data quantity, and reduce resource consumption.

[0093] In step 103, store the data features in the feature platform.

[0094] The feature platform is a platform with a feature storage function. This feature platform is, for example, a data storage system or a data storage device, and can be built based on software and / or hardware. In addition, the feature platform can also have functions such as data query, use, addition, and deletion to ensure the effective management of the data stored in the platform.

[0095] In this step, by introducing the feature platform, the data features extracted based on the data to be processed are stored, which is convenient for the subsequent effective retrieval and use of the data features, reduces the data storage volume, reduces the utilization complexity of the data to be processed with a large data volume and diverse and complex data forms, and reduces the data reading and calculation cost caused by directly accessing the data to be processed.

[0096] In the above data processing process, when performing feature extraction on the data to be processed, the data index type and the data feature level are introduced. Based on the data index type, multiple levels of data features are extracted, enabling the data to be processed to be represented in the form of multi-level data features in combination with the data index type, ensuring the diversity and multi-level nature of the features when representing the data to be processed in the form of data features, and being able to achieve a large-scale compression of the data to be processed, reducing the utilization complexity of the data to be processed with a large amount of data and diverse and complex data forms, reducing the data storage volume and the data reading and calculation costs, helping to improve the data processing efficiency, and reducing resource consumption.

[0097] On the one hand, in some embodiments, the implementation process of step 102 for performing feature extraction on the data to be processed according to the data index type to obtain data features at different levels may be as follows:

[0098] Directly based on multiple data index types, perform feature extraction on the data to be processed once according to each data index type, obtain a level of data features corresponding to each data index type respectively, and finally form multiple levels of data features. The data features at different levels respectively correspond to a data index type.

[0099] That is, perform feature extraction on the data to be processed independently according to each data index type to obtain data features at different levels.

[0100] Differently, in the embodiments of the present application, it is also set that there is a dependency relationship between different data index types. In this case, based on the dependency relationship between the data index types, perform layer-by-layer progressive feature processing on the data to be processed to obtain multiple levels of data features.

[0101] That is, on the other hand, in some embodiments, as shown in combination with Figure 3 Step 102 for performing feature extraction on the data to be processed according to the data index type to obtain data features at different levels includes:

[0102] Step 301, perform feature extraction on the data to be processed according to the first data index type to obtain the data features at the first level.

[0103] Step 302, according to the dependency relationship between different data index types, sequentially perform feature extraction of the subsequent data index type on the currently obtained data features at the current level to obtain the data features at the subsequent level until multiple levels of data features are obtained.

[0104] Among them, the dependency relationship between multiple data index types will change with the different selections of the data index type.

[0105] The dependency relationship can be a data dependency relationship in the time dimension or the population dimension. For example, the annual data indicator feature depends on the quarterly data indicator feature, and the quarterly data indicator feature depends on the monthly data indicator feature, forming a dependency relationship between data indicator types; the community indicator feature depends on the household indicator feature, and the household indicator feature depends on the individual indicator feature, etc. This is only an exemplary illustration and is not limited thereto.

[0106] In the above process, there are multiple data indicator types. The feature extraction results of the first data indicator type form the data features of the first level. After performing feature extraction once according to the first data indicator type, the feature extraction of the subsequent data indicator type is sequentially executed according to the dependency relationship between multiple data indicator types, and so on, until the data features of multiple levels are extracted. There is a dependency relationship between different data indicator types, which makes there also be a dependency relationship between the data features of multiple levels obtained by feature extraction. Through hierarchical feature processing, the data volume of different levels is compressed layer by layer, helping to improve the subsequent data calculation efficiency. Among them, in the process of performing the feature extraction of the subsequent data indicator type based on the data features of the currently obtained level, it can be to perform the feature extraction of the subsequent data indicator type only based on the data features of the current level, or to perform the feature extraction of the subsequent data indicator type by adding other data features while using the data features of the current level.

[0107] In an optional example, multiple data indicator types include, for example, block feature, event feature, persona feature, and business feature. Among them, the block feature specifically refers to the feature data after being segmented and aggregated according to specific rules. The event feature specifically refers to the feature data cut by different battery states. The persona feature specifically refers to the feature data borrowed from the concept of user persona, which can objectively reflect the operating conditions and state changes of entities. The business feature specifically refers to the feature data that can reflect certain business changes.

[0108] During the application process, the battery management system (BMS) in the vehicle collects data related to current or power, and then updates it at a frequency of once every 5 seconds or 1 second. A single charging process usually lasts for 50 - 70 minutes, and the data of a complete charging process usually contains 700 - 1000 pieces of data. Calculated according to the frequency of charging once every 3 days, there are about 100 chargings throughout the year. A passenger car has nearly 100,000 pieces of data related to charging, and a vehicle used for about 3 years will accumulate about 300,000 pieces of data.

[0109] Taking the data to be processed as the original data as an example, specifically taking "charging power" as the original data, and on this basis, according to several data index types such as block characteristics, event characteristics, portrait characteristics, and business characteristics, selecting the data content under each data index type from the original data to perform relevant data feature calculations as an example, the relevant features are shown in the following table:

[0110]

[0111] The dependency relationships among the several data index types of block characteristics, event characteristics, portrait characteristics, and business characteristics are as Figure 4 shown. The original data includes the current and voltage data of the battery. According to the data index type of block characteristics, the voltage and current data aggregated through a 5-minute time window are selected from the original data, and one feature processing is performed to obtain the charging power per 5 minutes; on this basis, according to the data index type of event characteristics, the charging power per 5 minutes is aggregated from it to obtain the charging power during each complete charging process, and one feature processing is performed to obtain the charging power during one complete charging process; on this basis, according to the data index type of portrait characteristics, the charging powers during multiple complete charging processes are aggregated from it to obtain the cumulative charging power since the battery was put into use, and one feature processing is performed to obtain the charging power charged during the entire life cycle of the battery; on this basis, according to the data index type of business characteristics, the average value of the annual charging amount per vehicle of a certain brand of vehicle is selected and aggregated from it, and one feature processing is performed to obtain the average annual charging amount of a certain vehicle model; finally, 4 levels of data features are obtained.

[0112] In this process, through the layer-by-layer compression of data features, the feature calculation of the latter level does not need to access the initial data to be processed. Taking the Figure 4 process in as an example, when calculating on data of 300,000 scale, using a 2-core 4G computing server, it takes about m minutes to calculate each type of feature, then the total resource consumption is 8m units.

[0113] The resource consumption of feature calculation gradually decreases as shown in the following table.

[0114] Data index type Data scale Resource consumption Original data 300,000 / Block characteristics 10,000 8m Event characteristics 100 8 / 30m Portrait characteristics 1 100 / 300000*8m Business characteristics 3 Negligible

[0115] In this process, after being reconstructed by feature engineering, the computing resource is reduced from 4 * 8m = 32m to 8.24m, greatly reducing the computing resource consumption.

[0116] In the above process, by using the dependency relationships among data index types, on the basis of this dependency relationship, the feature processing of the data to be processed is sequentially performed. Through the layer-by-layer progressive feature processing, different levels of data features are obtained. During the process of obtaining different levels of data features, the repeated access to the data to be processed is reduced, the feature processing efficiency is improved, the data reading and calculation cost is reduced, and the resource consumption is reduced.

[0117] In some embodiments, in combination with Figure 5 As shown, in step 302, feature extraction of the subsequent data metric type is performed on the data features of the current level obtained by extraction, to obtain data features of the subsequent level, including:

[0118] Step 501, aggregate the data features of the current level according to the target data granularity, to obtain aggregated data.

[0119] Among them, the data features of the current level extracted under the current data metric type have their data granularity. The target data granularity is greater than the data granularity of the data features of the current level.

[0120] For each level of data features obtained by extraction, the data features of this one level form the data features of the current level.

[0121] In an alternative embodiment, different data metric types correspond to different data granularities under the same data metric dimension.

[0122] In an example, the data metric dimension may be a time dimension, a region dimension, a population granularity, etc. The data granularity may be days, weeks, months under the time dimension; individual provinces, North China, South China, Northwest regions, etc. under the region dimension.

[0123] Step 502, perform feature extraction on the aggregated data according to the subsequent data metric type, to obtain data features of the subsequent level.

[0124] The subsequent data metric type refers to, among data metric types having a dependency relationship, the data metric type that depends on other data metric types to implement the feature processing of this data metric type forms the subsequent data metric type, and the corresponding feature processing operation needs to be executed after the feature processing of the data metric type being depended on is completed.

[0125] During the hierarchical extraction process of data features for data metric types having a dependency relationship, the feature extraction objects (being data to be processed or aggregated data) of different data metric types have different data granularities, and there are different feature granularities between the different levels of data features obtained by corresponding extraction.

[0126] During the hierarchical feature extraction process of aggregating feature content to form the feature extraction object corresponding to a certain level, the feature content in the data features of the current level can be integrated according to the size of the feature granularity. For example, multiple small-granularity feature contents in units of weeks are integrated into a data content set according to the large-granularity feature in units of months.

[0127] In this process, the generated data features are used as the data basis for generating the features of the subsequent data metric type. Based on the granularity differences between the data metric types with dependency relationships and the data features at different levels, the feature contents at different levels are aggregated, and feature extraction is performed on the aggregated data based on the subsequent data metric type to obtain the data features at the subsequent level, so as to realize the progressive processing of feature contents in a more efficient manner and improve the feature processing efficiency.

[0128] In some embodiments, in combination with Figure 6 As shown, after step 103 stores the data features in the feature platform, it further includes:

[0129] Step 601, query the first required features adapted to the algorithm engineering from the feature platform.

[0130] Step 602, when it is determined that there are first required features in the feature platform, input the first required features into the algorithm engineering to obtain the feature processing results output by the algorithm engineering.

[0131] The feature platform provides a data query function to ensure that after storing the data features, the algorithm engineering can retrieve and use the pre-generated data features, rather than directly using the data to be processed for evaluation and calculation.

[0132] In the above steps, the algorithm engineering corresponds to different algorithm models.

[0133] Specifically, after using feature engineering to extract more useful features from the data to be processed, these features are then used as the processing objects of the algorithm engineering, changing the practice of directly applying the algorithm engineering to the data to be processed. On the one hand, the data features can be used to optimize the machine learning algorithm to obtain the best model performance, facilitating the feature algorithm to better learn patterns and rules. On the other hand, the algorithm engineering can directly perform calculation processing on the data features to obtain evaluation results, realizing the sharing and common use of the data features in the feature platform and improving the information evaluation efficiency.

[0134] The above process changes the way of directly accessing the data to be processed, uses the different-level data feature extraction requirements stored in the feature platform to extract the required features for algorithm engineering utilization, realizes the sharing and common use of data features through the feature platform, improves the feature reuse rate, reduces the data processing complexity of evaluation and calculation, improves the data processing efficiency, and reduces resource consumption.

[0135] In some embodiments, after step 601 queries the first required features adapted to the algorithm engineering from the feature platform, it further includes:

[0136] In the case where there is no first required feature in the feature platform, obtain the selected data as the data to be processed, and return to execute step 102. According to the data index type, perform feature extraction on the data to be processed to obtain data features at different levels until there is a first required feature in the feature platform.

[0137] The query function in the feature platform can be used to query whether the required features for query calculation already exist in the platform. If they already exist, they can be directly used. If not, the required features need to be generated through the platform.

[0138] When selecting to generate the required new features, by selecting the data to be processed and performing feature processing on the selected data, new feature algorithms, new data index types, new data feature levels, etc. can be accompanied during the feature processing. After extracting the required features, store them in the feature platform so that the new features can be queried in the feature platform, realizing the expansion of the data features in the feature platform, enhancing the richness of the data features in the feature platform, and enhancing the effect of feature sharing and common use.

[0139] And this process is also applicable to the situation where there is no data stored in the feature platform. Through the above processing process, feature processing and storage are realized, and the quick black start of the feature platform is achieved.

[0140] In the above processing process, when querying the required features using the feature platform but not being able to query them, it can be determined that the selected data is used as the data to be processed, and data feature processing is performed on the selected data to generate data features at different levels, ensuring that the current required features are stored in the feature platform, realizing data expansion and system black start, and facilitating the effective development of subsequent information evaluation processing.

[0141] In some embodiments, after storing the data features in the feature platform in step 103, it further includes:

[0142] Select a second required feature from the feature platform and directly output the second required feature to the user side.

[0143] This process provides a feature output method, bypassing the information evaluation calculation process of algorithm engineering, and directly outputting the data features pre-processed in the feature platform to the user, meeting the feature output requirements of the data to be processed in specific situations, reducing the data evaluation processing flow, and expanding the application scenarios and information processing efficiency of the feature data in the feature platform.

[0144] In some embodiments, the data processing method further includes:

[0145] Select the target data content from the data to be processed and directly output the target data content to the user side.

[0146] This process provides a data output method that directly outputs appropriate data in the data to be processed to the user, meets the data output requirements in specific situations, and expands the data application scenarios and information processing efficiency.

[0147] In some embodiments, the data processing method further includes:

[0148] Evaluating the stored data features in the feature platform; when it is determined that the stored data features need to be iterated, perform step 102 to extract data features from the data to be processed according to the data index type, and obtain data features at different levels.

[0149] Among them, when the evaluation result of the stored data features does not meet the index requirements, it can be determined that these data features need to be iterated. The index requirements are, for example, that the storage duration is within the time threshold, the usage amount is not less than the number threshold, and it conforms to the professional knowledge in the data field, etc., and can be set according to the evaluation requirements.

[0150] This process ensures the timely iterative update of the data stored in the feature platform and the effectiveness of the features through the evaluation of the stored data features in the feature platform.

[0151] In some embodiments, as shown in Figure 7 The data processing method further includes:

[0152] Step 701: Evaluate the stored data features in the feature platform.

[0153] Among them, the evaluation of the stored data features is, for example, to evaluate whether the storage duration of the data features exceeds the time threshold, whether the usage amount of the data features is lower than the number threshold, or to evaluate the effectiveness of the data features in combination with the professional knowledge in the data field, etc., and can be set according to the evaluation requirements.

[0154] Step 702: When it is determined that there are target data features to be taken offline in the stored data features, take the target data features offline from the feature platform.

[0155] When the evaluation result does not meet the index requirements, it can be determined that these data features are outdated or not frequently used, and they can be taken offline.

[0156] This process ensures the timely offline update of the data stored in the feature platform and the effectiveness of the features through the evaluation of the stored data features in the feature platform.

[0157] In the above embodiments, through the constructed feature platform, querying of existing data features stored in the platform, rapid generation of data features, invocation of existing data features in the platform, timely updating or taking offline of infrequently used features are realized. Through this feature platform, closed-loop management of the entire life cycle of data features can be achieved, development, sharing, and management of data features can be realized, and feature sharing among different evaluation algorithm projects can be satisfied, reducing the calculation cost, providing room for further improvement of the model calculation accuracy, and increasing the reuse rate of features.

[0158] In some embodiments, before step 102 extracts features from the data to be processed according to the data index type to obtain data features at different levels, it further includes:

[0159] Configuring a feature algorithm and task scheduling for the feature algorithm for the data to be processed.

[0160] According to the feature algorithm and task scheduling, step 102 is executed to extract features from the data to be processed according to the data index type to obtain data features at different levels.

[0161] Among them, the data feature levels are indicated in the feature algorithm.

[0162] When generating new features, it is first necessary to determine the data to be processed. Based on the determined data to be processed, a feature calculation method is configured. Existing algorithms can be selected through the feature platform, or special professional algorithms can be registered with the feature platform for selection and use.

[0163] Among them, the feature algorithm corresponds to data feature processing. In one implementation process, the feature algorithm corresponds to the data index type used when performing data feature processing based on the data to be processed. The two can be in a one-to-many relationship or a many-to-one relationship, which can be set as needed and is not limited thereto here.

[0164] At the same time, task scheduling also needs to be configured. The execution period of the feature calculation task can be configured according to the calculation period requirements of the data to be processed and the feature algorithm, such as real-time, minute, hour, day, week, etc.

[0165] The configuration of task scheduling can be that after the feature platform obtains the configuration information of the feature algorithm, it automatically coordinates and configures the task scheduling according to the calculation needs, or it can be directly set by the user.

[0166] Subsequently, according to the feature algorithm and task scheduling, the process of extracting features from the data to be processed according to the data index type to obtain data features at different levels is implemented, and finally the features are stored in the feature platform.

[0167] This process ensures the reasonable scheduling of computing resources and improves the feature calculation efficiency by configuring the feature algorithm and the corresponding task scheduling.

[0168] In some embodiments, before configuring a feature algorithm and task scheduling of the feature algorithm for data to be processed, the following steps are further included:

[0169] Obtain the feature algorithm selected by the user and the feature spanning period; determine the task scheduling corresponding to the feature algorithm based on the feature spanning period.

[0170] The feature spanning period, such as real-time, minute, hour, day, week, month, or even multi-time scales of the entire life cycle, can be selected and configured by the user to meet the feature processing requirements.

[0171] When determining the task scheduling corresponding to the feature algorithm based on the feature spanning period, the scheduling implementation plan of tasks involved in the feature algorithm, such as data screening, data aggregation, and data feature extraction, can be directly set according to the time scale of the feature spanning period, that is, the task scheduling is determined.

[0172] This process enables the user to independently select the current required feature calculation method, and determine the task scheduling based on the feature algorithm selected by the user and the corresponding feature spanning period. While meeting the actual needs of the user, it ensures the reasonable scheduling of computing resources and improves the feature calculation efficiency.

[0173] Next, in combination with the specific application process of battery data, the implementation process of the data processing method will be described as a whole by way of example.

[0174] Combined with Figure 8 As shown, obtain the battery data uploaded by each vehicle to form raw data. In this case, the data to be processed in the embodiments of the present application is this raw data. Feature processing needs to be performed on these raw data to obtain battery data features at multiple levels.

[0175] Specifically, in the processing process, the raw data enters the feature engineering for feature processing. Through feature engineering, feature extraction at different levels can be achieved, generating various types of data features including block features, event features, battery portraits, and business features.

[0176] Store the generated data in the feature platform, and the data in the feature platform can be shared and used in common.

[0177] In the subsequent process, the data features extracted by the feature engineering enter the algorithm engineering for use by the algorithm model. During the use process, different algorithm models in the algorithm engineering can all call the data features output by the feature engineering through the feature platform to achieve the reuse of data features in multiple algorithms.

[0178] The algorithm model in the algorithm engineering can call the data features to generate battery alarms or other processing results, and release them to the user side through the user engineering for use.

[0179] In addition, the data to be processed as the original data can still be directly connected to the user project for data output and consumed by the user side. On the premise of minimizing the number of access and usage times of the data to be processed, the original data can still be directly connected to the algorithm project for its call and consumption. The feature data output by the feature engineering and stored in the feature platform can be directly connected to the user project for data output and consumed by the user side, especially the battery portrait data among them.

[0180] In the above entire process, the introduction of feature engineering and the feature platform enables the data to be processed to be characterized in the form of multi-level data features, greatly compressing the data volume of the data to be processed, reducing the utilization complexity of the data to be processed with a large data volume and diverse and complex data forms, realizing the closed-loop management of the entire life cycle of data features, realizing the sharing and common use of feature data, improving data processing efficiency, and reducing resource consumption.

[0181] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0182] Based on the same inventive concept, the embodiments of the present application also provide a data processing device. The data processing device provided by the embodiments of the present application can implement each process of the embodiments of the above data processing method and can achieve the same technical effects. Therefore, the specific limitations in one or more of the following data processing device embodiments can refer to the limitations on the data processing method in the above text. To avoid repetition, they will not be elaborated here.

[0183] In one embodiment, as Figure 9 shown, a data processing device 900 is provided, including:

[0184] A data acquisition module 901, configured to acquire the data to be processed and the data index type;

[0185] A feature generation module 902, configured to extract features from the data to be processed according to the data index type to obtain data features at different levels;

[0186] A feature storage module 903 for storing the data features into a feature platform.

[0187] In some embodiments, the feature generation module 902 is specifically configured to:

[0188] Extract features from the data to be processed according to the first data metric type to obtain data features at the first level;

[0189] According to the dependency relationship between different data metric types, sequentially perform feature extraction of the subsequent data metric type on the currently extracted data features at the current level to obtain data features at the subsequent level until multiple levels of data features are extracted.

[0190] In some embodiments, the feature generation module 902 is more specifically configured to:

[0191] Aggregate the data features at the current level according to the target data granularity to obtain aggregated data; the target data granularity is greater than the data granularity of the data features at the current level;

[0192] Extract features from the aggregated data according to the subsequent data metric type to obtain data features at the subsequent level.

[0193] In some embodiments, the apparatus further includes:

[0194] An algorithm engineering module for:

[0195] Query the first required features adapted to the algorithm engineering from the feature platform;

[0196] In the case where it is determined that the first required features exist in the feature platform, input the first required features into the algorithm engineering to obtain the feature processing result output by the algorithm engineering.

[0197] In some embodiments, the apparatus further includes:

[0198] A data augmentation module for:

[0199] In the case where it is determined that the first required features do not exist in the feature platform, obtain selected data as the data to be processed, and return to execute the step of extracting features from the data to be processed according to the data metric type to obtain data features at different levels until it is determined that the first required features exist in the feature platform.

[0200] In some embodiments, the apparatus further includes:

[0201] A first feature output module for selecting second required features from the feature platform and directly outputting the second required features to the user terminal.

[0202] In some embodiments, the apparatus further includes:

[0203] A second feature output module, configured to select target data content from the data to be processed and directly output the target data content to the user side.

[0204] In some embodiments, the apparatus further includes:

[0205] A data iteration module, configured to:

[0206] Evaluate the stored data features in the feature platform;

[0207] In the case where it is determined that the stored data features need to be iterated, perform the step of extracting data features of different levels from the data to be processed according to the data metric type.

[0208] In some embodiments, the apparatus further includes:

[0209] A data offline module, configured to:

[0210] Evaluate the stored data features in the feature platform;

[0211] In the case where it is determined that there are target data features to be offline processed in the stored data features, offline the target data features from the feature platform.

[0212] In some embodiments, the apparatus further includes:

[0213] A configuration module, configured to:

[0214] Configure a feature algorithm and task scheduling of the feature algorithm for the data to be processed, where N levels are indicated in the feature algorithm; perform the step of extracting data features of different levels from the data to be processed according to the data metric type according to the feature algorithm and the task scheduling.

[0215] In some embodiments, the configuration module is further configured to:

[0216] Obtain the feature algorithm and feature spanning period selected by the user;

[0217] Determine the task scheduling corresponding to the feature algorithm based on the feature spanning period.

[0218] Each module in the above data processing apparatus can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0219] In one embodiment, as Figure 10 shown, a computer device is provided. The computer device 10 of this embodiment includes: at least one processor 1000 ( Figure 10 only one is shown in the figure), a memory 1001, and a computer program 1002 stored in the memory 1001 and executable on the at least one processor 1000. When the processor 1000 executes the computer program 1002, the steps in any of the above method embodiments are implemented.

[0220] The computer device 10 may be a computing device such as a desktop computer, a notebook, a palm computer, etc. The computer device 10 may include, but is not limited to, a processor 1000 and a memory 1001. Those skilled in the art can understand that Figure 10 these are merely examples of the computer device 10 and do not constitute a limitation on the computer device 10. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.

[0221] The processor 1000 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0222] The memory 1001 may be an internal storage unit of the computer device 10, such as the hard disk or memory of the computer device 10. The memory 1001 may also be an external storage device of the computer device 10, such as a plug-in hard disk equipped on the computer device 10, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 1001 may also include both the internal storage unit and the external storage device of the computer device 10. The memory 1001 is used to store the computer program and other programs and data required by the computer device. The memory 1001 may also be used to temporarily store data that has been output or will be output.

[0223] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0224] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0225] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0226] In the embodiments provided in this application, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0227] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0228] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0229] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-mentioned embodiment methods of the present application may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0230] All or part of the processes in the above-mentioned embodiment methods of the present application may also be implemented through a computer program product. When the computer program product runs on a computer device, the computer device is caused to execute the steps in the above-mentioned various method embodiments.

[0231] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A data processing method, characterized in that, it includes: obtaining data to be processed and data index types; extracting features from the data to be processed according to the data index types to obtain data features at different levels; storing the data features in a feature platform.

2. The method according to claim 1, characterized in that, the extracting features from the data to be processed according to the data index types to obtain data features at different levels includes: extracting features from the data to be processed according to a first data index type to obtain data features at a first level; successively extracting features of a subsequent data index type from the currently extracted data features according to the dependency relationship between different data index types until data features at multiple levels are extracted.

3. The method according to claim 2, characterized in that, the extracting features of a subsequent data index type from the currently extracted data features to obtain data features at a subsequent level includes: aggregating the currently extracted data features according to a target data granularity to obtain aggregated data; the target data granularity is greater than the data granularity of the currently extracted data features; extracting features from the aggregated data according to the subsequent data index type to obtain data features at a subsequent level.

4. The method according to any one of claims 1 to 3, characterized in that, after storing the data features in the feature platform, it further includes: querying a first required feature adapted to an algorithm project from the feature platform; when it is determined that the first required feature exists in the feature platform, inputting the first required feature into the algorithm project to obtain a feature processing result output by the algorithm project.

5. The method according to claim 4, characterized in that, after querying the first required feature adapted to the algorithm project from the feature platform, it further includes: when it is determined that the first required feature does not exist in the feature platform, obtaining selected data as the data to be processed, and returning to execute the step of extracting features from the data to be processed according to the data index types to obtain data features at different levels until it is determined that the first required feature exists in the feature platform.

6. The method according to any one of claims 1 to 5, characterized in that, after storing the data features in the feature platform, it further includes: selecting a second required feature from the feature platform and directly outputting the second required feature to a user terminal.

7. The method according to any one of claims 1 to 6, characterized in that, it further includes: selecting target data content from the data to be processed and directly outputting the target data content to a user terminal.

8. The method according to any one of claims 1 to 7, characterized in that, it further includes: evaluating the stored data features in the feature platform; when it is determined that the stored data features need to be iterated, executing the step of extracting features from the data to be processed according to the data index types to obtain data features at different levels.

9. The method according to any one of claims 1 to 8, wherein, further comprising: evaluating the stored data features in the feature platform; when it is determined that there are target data features to be taken offline in the stored data features, taking offline the target data features from the feature platform.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer program product, wherein, comprising computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in an electronic device, the processor in the electronic device executes the steps of the method according to any one of claims 1 to 9.