Automobile sales volume multi-dimensional prediction model construction system and method
The multi-dimensional prediction model is obtained through the data acquisition unit, the dynamic feature engineering unit generates multi-dimensional prediction features, the multi-model collaborative prediction unit constructs a time series-non-time series dual-channel prediction framework, the model training optimization unit optimizes the sub-model parameters, and the prediction result output unit generates sales forecast values and confidence intervals. This solves the problems of insufficient dynamic adaptability and fusion adaptability in existing technologies and improves prediction accuracy and stability.
Patent Information
- Application Number
- CN202511284892.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack dynamic adaptability in automobile sales forecasting, feature processing is not flexible enough, the prediction framework and result fusion are not adaptable enough, and data preprocessing and model training are not refined enough, resulting in poor prediction accuracy and stability.
Build a multi-dimensional prediction model for automobile sales, obtain multi-dimensional data through the data acquisition unit, generate multi-dimensional prediction features through the dynamic feature engineering unit, build a time series-non-time series dual-channel prediction framework through the multi-model collaborative prediction unit, optimize the sub-model parameters through the model training optimization unit, and generate sales forecast values and confidence intervals through the prediction result output unit.
It realizes dynamic adaptive adjustment of the prediction feature set, balances the prediction contribution of time series and non-time series dimension features, improves prediction accuracy and stability, provides a more reliable data foundation and result display, and supports enterprises in providing effective data for production scheduling and decision-making.
Smart Images

Figure CN120765302A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automobile sales prediction, in particular to a multi-dimensional automobile sales prediction model construction system and method. BACKGROUND
[0002] Automobile sales prediction is a key technical support for automobile enterprise production scheduling, inventory control and market strategy optimization. With the advancement of digitalization in the automobile industry, the influencing factors of sales have expanded from traditional market demand and price factors to multi-dimensional data such as vehicle production records, user operation trajectories, and regional environmental monitoring. These multi-source heterogeneous data have the characteristics of dynamic change and complex correlation. Traditional prediction methods that rely on a single data source or static models cannot effectively capture the dynamic characteristics of data and the coupling effects of multiple factors, resulting in insufficient prediction accuracy and poor adaptability, which cannot meet the fine-grained needs of enterprises for sales prediction.
[0003] In the prior art, related patents have carried out research on automobile sales prediction through data preprocessing and model optimization. For example, Chinese patent CN202410518245.7 discloses a new energy automobile sales prediction method, system and medium based on a fusion model, which includes: obtaining new energy automobile data, completing the data through linear interpolation strategy; using the isolated forest method to remove abnormal data to obtain initial experimental data; extracting specific data corresponding to the first automobile key data of the pre-set key data features; converting it into second automobile key data of a pre-set time granularity; training a pre-set fusion model based on the second automobile key data, and obtaining a sales prediction model after training. For another example, Chinese patent CN202111356919.0 discloses a multi-scale information automobile sales big data prediction method based on an attention mechanism, which includes: transmitting automobile basic information data with user behavior information into an encoderRNN; obtaining multi-scale information of each time sequence at each time step in the encoderRNN by using multi-scale feature decoupling operation; obtaining the importance score of different smallhiddenstate by using the attention mechanism and updating; outputting the automobile sales prediction result of the future time step by the decoderRNN to dynamically select important scale information to improve the prediction accuracy.
[0004] Although the above technical solution has corresponding design advantages, it still has the following technical defects: First, the feature processing lacks dynamic adaptability: Chinese patent CN202410518245.7 relies on a preset set of key data features to carry out predictions, without considering the changes in the effectiveness of features in different prediction cycles, and cannot actively eliminate long-term inefficient features or incorporate newly generated efficient features; Although Chinese patent CN202111356919.0 obtains time series information through multi-scale feature decoupling, it does not combine the influence and real-time contribution of features on the prediction results for dynamic screening, resulting in the adaptability of features to prediction requirements decreasing with changes in scenarios; Second, the adaptability of the prediction framework and result fusion is insufficient: Chinese patent CN202410518245.7 does not distinguish between time series features (such as historical sales trends) and non-time series features (such as product technology). The Chinese patent CN202111356919.0 focuses solely on time-series data processing, ignoring the coupled impact of non-time-series features such as product configuration differences and regional environmental monitoring on sales. Furthermore, neither method dynamically adjusts the fusion weights of the model output based on real-time prediction deviations, making it difficult to balance the prediction contributions of different data dimensions and impacting the stability of prediction accuracy. Third, data preprocessing and model training are insufficiently refined: Chinese patent CN202410518245.7 fails to design differentiated missing value imputation strategies for continuous and discrete data, and outlier processing utilizes only the isolation forest method without optimizing data distribution characteristics. Chinese patent CN202111356919.0 also fails to avoid time periods with concentrated data mutations during model training, which can easily reduce the generalization ability of the trained model due to local data distribution shifts. Therefore, we propose a system and method for constructing a multi-dimensional prediction model for automobile sales. Summary of the Invention
[0005] The purpose of the present invention is to provide a system and method for constructing a multi-dimensional prediction model for automobile sales, so as to solve the problems raised in the above background technology.
[0006] To solve the above technical problems, one of the objectives of the present invention is to provide a system for constructing a multi-dimensional prediction model for automobile sales, comprising: A data acquisition unit, which is used to obtain multi-dimensional vehicle-related raw data. By connecting to the vehicle production database, sales terminal recording system, user behavior sensing equipment and government data disclosure interface, the unit collects and stores vehicle production records, historical sales data, product technical parameters, user operation trajectories and environmental monitoring data; A data preprocessing unit, which is used to clean and standardize the raw data, fill missing values using statistical interpolation, identify and correct outliers based on data distribution characteristics, and convert unstructured data into standardized data in a unified format using a numerical conversion algorithm; A dynamic feature engineering unit, which is used to generate a multi-dimensional prediction feature set and adaptively extract key features through a dynamic screening mechanism that quantifies feature importance and provides real-time contribution feedback. The dynamic screening mechanism is based on the calculation of the impact of features on prediction error and combines a sliding window to iteratively update the feature set; A multi-model collaborative prediction unit is used to build a time series and non-time series dual-channel prediction framework. It uses a recurrent neural network with an attention mechanism to process time series features, a tree-structured ensemble learning algorithm to process product attributes and user behavior characteristics, and an adaptive weight adjustment algorithm based on real-time prediction deviations to achieve multi-sub-model output fusion. A model training and optimization unit, which is used to perform parameter training and performance optimization on the sub-models of the multi-model collaborative prediction unit. The training set, validation set, and test set are divided based on the historical timestamps corresponding to the prediction features output by the dynamic feature engineering unit. The core parameters of the sub-models are iteratively updated using a gradient descent optimizer. Training is terminated when the validation set error does not improve for consecutive preset rounds. If overfitting occurs or the error does not improve, the sub-model parameters are adjusted or some data is supplemented before restarting training. A prediction result output unit is used to generate and output automobile sales forecast results. By combining the trained model parameters output by the model training optimization unit and the real-time features of the dynamic feature engineering unit, it generates sales forecast values and confidence intervals for a future preset period. The forecast values, confidence intervals and core feature influence weights are presented through a visual interface. The structured export of prediction results is supported, and historical sales data stored in the data acquisition unit can be called to achieve a retrospective comparison between the prediction results and actual historical sales.
[0007] As a further improvement of this technical solution, the data acquisition unit includes an interface protocol adaptation module, an edge data preprocessing module, a classification storage module, a collection scheduling module and an integrity verification module, wherein: The interface protocol adapter module is used to achieve standardized access to multi-source data. By integrating the MQTT protocol adapter component, HTTP protocol adapter component, HTTPS protocol adapter component, and OPCUA protocol adapter component, it connects to the automobile production database, sales terminal recording system, user behavior sensor equipment, and government data disclosure interface respectively, and converts the raw data output by heterogeneous data sources into a unified format that can be recognized by the system. The edge data preprocessing module is used to process the high-frequency raw data output by the user behavior sensor device, using the LZ4 compression algorithm to reduce the data transmission bandwidth occupancy and the sliding window filtering algorithm to filter out the high-frequency noise during the collection process; The classified storage module is used for storing by data characteristics, and specifically comprises: vehicle production records and product technical parameters are stored in a relational database, taking production batch number as a unique index; historical sales data and user operation track are stored in a time series database, taking "time stamp+region code" as a composite index; and environmental monitoring data are stored in a distributed file system, taking the longitude and latitude grid coordinates of the monitoring area as a storage path. The collection scheduling module is used for dynamically configuring a collection strategy: static data adopts daily fixed time period full quantity synchronization; dynamic data adopts change log triggered incremental collection; and environmental monitoring data set a timing polling period according to index types (meteorological data 1 hour each time, and traffic flow data 15 minutes each time). The integrity checking module is used for verifying the integrity of data transmission and storage, and performs transmission verification by calculating the SHA-256 hash value of the data, and performs storage verification by key field non-empty verification and data number consistency verification, and triggers a retransmission mechanism when the verification fails.
[0008] As a further improvement of the technical solution, the data preprocessing unit comprises a missing value filling module, an abnormal value identification and correction module, a data conversion module and a preprocessing verification module, wherein: The missing value filling module is used for processing missing items in original data, and specifically comprises: for continuous numerical data, a linear interpolation method is adopted to calculate a filling value based on the trend of effective data points before and after the missing position; and for discrete classification data, a weighted mode method is adopted, and the category with the highest frequency in the adjacent same data of the missing item is taken as the filling value, and the frequency weight of recent data is higher than that of long-term data. The abnormal value identification and correction module is used for detecting and correcting values deviating from data distribution characteristics, and specifically comprises: for numerical data subject to normal distribution, whether the deviation of the data point from the mean value exceeds 3 times the standard deviation is judged to identify abnormal values; for non-normal distribution data, a quartile range method is adopted; and the identified abnormal values are corrected by a local weighted regression method, and a correction value is fitted based on the trend of multiple effective data points before and after the abnormal value. The data conversion module is used for realizing data standardization and structured conversion, and specifically comprises: for numerical data, a min-max standardization is adopted to linearly map the data to a preset interval; for discrete type data, a one-hot encoding is adopted to convert each category into a binary feature vector; and for unstructured data, a bag-of-words model is adopted to extract keyword features and convert them into a structured frequency matrix. The preprocessing verification module is used for verifying the effectiveness of the processed data, and specifically comprises: whether the missing values are completely filled, whether the abnormal values conform to the overall data distribution trend after correction, and whether the standardized data are in the preset interval are checked for verification, and when any condition is not met, secondary processing of the corresponding module is triggered until all data meet the preprocessing requirements.
[0009] As a further improvement of the present technical solution, the dynamic feature engineering unit includes a multi-dimensional feature generation module, which includes a time feature extraction submodule, a product feature extraction submodule, a user feature extraction submodule, and an environment feature extraction submodule, wherein: The time feature extraction submodule is used to extract prediction features from time series data and extract periodic features through the time series period decomposition algorithm. , based on the quantification of the repeated fluctuation pattern of data within a fixed period; extracting trend features through sliding window linear fitting method , based on the change slope of the data in the window to reflect the long-term change direction; the mutation characteristics are extracted by the adjacent window difference threshold method , identify significant time nodes that deviate from the trend; The product feature extraction submodule is used to extract prediction features from technical parameters and extract configuration difference features by calculating the Euclidean distance of parameter vectors. , quantify the differences in hardware parameters of different models; extract performance matching features through parameter combination synergy analysis ,The association relationship based on key parameters reflects the ,adaptability of the product to the usage scenario; The user feature extraction submodule is used to extract prediction features from the operation trajectory and extract preference stability features by calculating the entropy value of the operation sequence. , quantify the consistency of users' long-term behavioral habits; extract behavioral conversion characteristics through key path conversion rate analysis , based on the user's operation chain from browsing to ordering, it reflects the stage characteristics of the decision-making process; The environmental feature extraction submodule is used to extract prediction features from environmental data and generate regional correlation features by mining the co-occurrence patterns of meteorological, traffic data and historical sales. , quantify the impact of environmental factors on sales in different regions.
[0010] As a further improvement of the present technical solution, the dynamic feature engineering unit further includes a dynamic screening mechanism module, which includes a feature importance quantification submodule, a real-time contribution feedback submodule, a window iterative update submodule, and a feature set stability verification submodule, wherein: The feature importance quantification submodule is used to calculate the influence of the prediction features extracted by the multi-dimensional feature generation module on the prediction results. ,It is determined by comparing the difference in prediction error after including the prediction feature and randomly replacing the prediction feature value.,The influence value is positively correlated with the error change amplitude; The real-time contribution feedback submodule is used to monitor the actual role of various prediction features in recent predictions. , based on the weighted cumulative value of feature importance within a sliding window (the window length is dynamically adjusted according to the prediction period). The weight coefficient of recent data in the window is higher than that of long-term data to ensure the timeliness of feedback; The window iteration update submodule is used to dynamically adjust the feature set. When the influence of any prediction feature extracted by the multi-dimensional feature generation module is Two consecutive windows are below the preset threshold, or their contribution When three consecutive windows show a monotonically decreasing trend, the prediction feature is automatically removed from the current feature set; at the same time, a new prediction feature is generated by combining cross-dimensional features. When the initial influence of the new prediction feature is When it exceeds the average level of the current feature set, it is included in the feature set to form an updated feature set; The feature set stability check submodule is used to verify the validity of the updated feature set by calculating the overlap rate between the current feature set before the update and the feature set generated after the update by the window iteration update submodule. and the forecast error fluctuation value To evaluate, when the overlap ratio Lower than the preset ratio or error fluctuation value When the threshold is exceeded, the multi-dimensional feature generation module is triggered to re-extract prediction features to optimize the feature set.
[0011] As a further improvement of the present technical solution, the multi-model collaborative prediction unit includes a time series feature processing module and a non-time series feature processing module, wherein: The time series feature processing module is used to process time series features, including a recurrent neural network submodule and an attention mechanism submodule, wherein: The recurrent neural network submodule adopts the long short-term memory network structure to extract the periodic features generated by the time feature submodule. , trend characteristics and mutation characteristics It takes the input as input, captures the dependencies of different time scales through the gating mechanism of input gate, forget gate and output gate, and outputs the hidden layer representation sequence of temporal features; The attention mechanism submodule is used to enhance the influence weight of key temporal features and calculate the attention weight of each time step based on the hidden layer representation sequence. , for mutation features The time step and the recent time step in the prediction window are given higher weights, and the hidden layer representation sequence is converted into an aggregate vector of time series features through weighted summation as the intermediate prediction result of the time series channel; The non-temporal feature processing module is used to process product attributes and user behavior features, including a tree-structured ensemble learning submodule and a feature interaction submodule, wherein: The tree ensemble learning submodule adopts the gradient boosting tree algorithm to extract the configuration difference features generated by the product feature extraction submodule. , performance matching characteristics , and preference stable features generated by the user feature extraction submodule , behavioral conversion characteristics As input, a nonlinear prediction model is constructed through iterative training of multiple decision trees, and the basic prediction value of non-time series features is output; The feature interaction submodule is used to mine cross-dimensional feature associations and to extract regional association features generated by the environmental feature extraction submodule. As a link, calculate the interaction strength between product features and user features , generate high-order interaction features, and input the high-order interaction features into the tree-structured ensemble learning submodule for secondary training, and correct the basic prediction value to form the final output of the non-time series channel.
[0012] As a further improvement of the present technical solution, the multi-model collaborative prediction unit further includes a model fusion module, which includes an adaptive weight adjustment submodule, a prediction deviation monitoring submodule and a fusion result verification submodule, wherein: The adaptive weight adjustment submodule is used to dynamically fuse the output results of the time series feature processing module and the non-time series feature processing module, and calculate the weight coefficient based on the real-time prediction deviation. : Time series model weight and non-series model weights The sum is 1, the weight value is generated by the inverse deviation function, and the weight update cycle is consistent with the prediction cycle; The prediction deviation monitoring submodule is used to quantify the submodule output error and calculate the timing deviation of the intermediate prediction results of the timing feature processing module based on the actual sales data. (square difference between the predicted value and the actual value) and the non-time series deviation finally output by the non-time series feature processing module , the deviation sequence is stored through a sliding window with a length of 5 prediction periods to provide a historical basis for weight adjustment; The fusion result verification submodule is used to verify the reliability of the fusion prediction by calculating the mean absolute error between the fusion result and the actual sales volume. And the improvement rate of the fusion result compared with the output of the time series feature processing module and the non-time series feature processing module Assessment: When Exceeds the preset threshold or When , the adaptive weight adjustment submodule is triggered to recalculate the weights, and the deviation-sensitive features are fed back to the time series feature processing module and the non-time series feature processing module to drive the secondary optimization of the feature weights.
[0013] As a further improvement of this technical solution, the model training optimization unit includes a data set time series partitioning module, a sub-model parameter iteration module, a training termination and adjustment module, and a training effect verification module, wherein: The dataset time series partitioning module is used to partition the training data according to the time series characteristics. Based on the historical timestamps corresponding to the prediction features output by the dynamic feature engineering unit, the dataset is divided into a training set, a validation set, and a test set in chronological order. When partitioning, ensure that the training set covers at least 2 complete The corresponding historical data, validation set, and test set each cover at least 1 complete Corresponding historical data, while avoiding mutation characteristics Concentrated time intervals to avoid data distribution shifts that affect training stability; The sub-model parameter iteration module is used to execute the sub-model core parameter update, and the gradient descent optimizer is used to optimize the two types of sub-models of the multi-model collaborative prediction unit respectively: for the time series feature processing module with attention mechanism recurrent neural network, its network connection weights and attention weights are optimized. Calculate coefficients and recurrent layer gating biases; optimize the decision tree feature splitting threshold and node splitting gain coefficient for the tree-structured ensemble learning model of the non-time-series feature processing module; dynamically set the initial learning rate based on the complexity of the sub-model, and decay at a fixed rate for each preset training round. During the decay process, the error changes of the validation set are recorded simultaneously; The training termination and adjustment module is used to trigger training termination and handle abnormal training status: the preset termination condition is "the decrease in the validation set error for 3-5 consecutive rounds is lower than the preset error threshold", and training is terminated when the condition is met; if the overfitting state of "the training set error continues to decrease while the validation set error increases" occurs, the time series sub-model is adjusted by increasing the dropout probability and the non-time series sub-model is adjusted by increasing the tree pruning intensity; if the errors of the training set and validation set do not improve for multiple consecutive rounds, less than or equal to 20% of recent data are selected from the test set to supplement the training set, and the parameter iteration process is restarted; The training effect verification module is used to verify the effectiveness of the trained sub-model, taking the test set that does not participate in parameter iteration as input and calculating the mean absolute percentage error between the predicted value and the actual sales volume. , goodness of fit ;when and When all the preset model performance thresholds are met, the training is deemed effective and the final sub-model parameters are output; if the threshold requirements are not met, the module returns to the dataset time series partitioning module to readjust the partitioning ratio and execute the training process again.
[0014] As a further improvement of this technical solution, the prediction result output unit includes a prediction result generation module, a visualization presentation module, a structured export module and a historical backtracking comparison module, wherein: The prediction result generation module is used to generate sales forecast data, receive the trained model parameters output by the model training optimization unit, load the real-time features output by the dynamic feature engineering unit, calculate and output the sales forecast value for the corresponding period according to the preset prediction period, and calculate the confidence interval corresponding to the prediction value based on the model prediction error characteristics; The visualization module is used to display forecast-related information, and simultaneously presents sales forecast values, corresponding confidence intervals, and the impact weights of core forecast features through the interface. The forecast values and confidence intervals are displayed in the form of trend charts, and the impact weights of core features are quantitatively presented in simple charts. The structured export module is used to output the forecast result file, supporting the export of sales forecast values, confidence intervals, and core feature impact weights in a standardized structured format. During the export process, the integrity of key data fields is automatically verified to ensure that there are no missing data. The historical backtracking comparison module is used to compare the forecast with historical data, call the historical sales data stored in the data acquisition unit, align the forecast results with the historical actual sales data according to the time dimension corresponding to the forecast period, calculate the deviation value between the two, and present the backtracking results in the form of a table or comparison chart.
[0015] A second object of the present invention is to provide a method for constructing a multi-dimensional prediction model for automobile sales, based on the above-mentioned multi-dimensional prediction model construction system for automobile sales, comprising the following steps: S100, Multi-source Automotive Raw Data Collection: Connect to automotive production databases, point-of-sale recording systems, user behavior sensing devices, and government data disclosure interfaces through interface protocol adaptation; process high-frequency raw data from user behavior sensing devices through edge data preprocessing; store data in classified storage partitions based on data characteristics; dynamically configure collection strategies through collection scheduling; verify data transmission and storage integrity through integrity checks, with verification failures triggering retransmission; S200, Data Cleaning and Standardization: Missing items in the original data are processed through missing value filling. Continuous numerical data uses linear interpolation, and discrete categorical data uses the weighted majority method. Outliers are detected and corrected through outlier identification and correction. Normally distributed data is identified based on 3 times the standard deviation, and non-normally distributed data is identified using the interquartile range method. Outliers are corrected using local weighted regression. Data standardization and structural transformation are achieved through data conversion. Numerical data uses min-max standardization, categorical data uses one-hot encoding, and unstructured data uses the bag-of-words model. The validity of the processed data is verified through preprocessing verification. S300, Dynamic Feature Set Generation: Extract multi-dimensional prediction features through multi-dimensional feature generation; screen features through a dynamic screening mechanism, calculate feature influence and real-time contribution, and iteratively update the feature set; verify the validity of the updated feature set through feature set stability check, and re-extract features if it does not meet the requirements; S400, Dual-channel Collaborative Prediction and Fusion: Performs temporal feature processing on time series features, employs recurrent neural networks to capture time-scale dependencies, and incorporates attention mechanisms to enhance the impact of key temporal features. It also performs non-temporal feature processing on product attributes and user behavior features, constructs a prediction model using tree-structured ensemble learning, and mines cross-dimensional feature associations through feature interaction. It also fuses the outputs of the two channels, calculates adaptive weights, monitors prediction deviations, verifies the fusion results, and optimizes them. S500, Model Training Optimization: The training set, validation set, and test set are divided into time series by data set time series to ensure that the complete cycle data is covered and the intervals with concentrated mutation features are avoided; the sub-model parameters are updated by gradient descent optimizer through sub-model parameter iteration, and the initial learning rate is dynamically set and decayed by rounds; training termination is triggered by training termination and adjustment to deal with overfitting and training stagnation; the effectiveness of the sub-model after training is verified by training effect verification. If it does not meet the requirements, the data set division is readjusted and training is resumed; S600, forecast result output: Generate sales forecast value and corresponding confidence interval based on forecast results; display forecast value, confidence interval and core feature influence weight through visualization; export forecast results in standardized format through structured export and verify data integrity; align forecast results with historical sales data through historical backtracking comparison, calculate deviation and present comparison results.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention uses a dynamic feature engineering unit to achieve dynamic screening and iterative updating of multi-dimensional prediction features. Combined with feature importance quantification, real-time contribution feedback, and feature set stability verification, it enables the prediction feature set to adaptively adjust with the prediction cycle and data changes, effectively preventing inefficient features from interfering with the prediction process and improving the adaptability of features to sales forecasting requirements. 2. This invention builds a dual-channel forecasting framework for time series and non-time series, relying on a multi-model collaborative forecasting unit. It uses a recurrent neural network with an attention mechanism to process time series features and a tree-structured ensemble learning algorithm to process non-time series features. It also integrates the outputs of multiple sub-models through adaptive weight adjustment based on real-time forecast deviations. This balances the forecast contributions of time series and non-time series features, reduces the impact of the limitations of a single forecast dimension on the results, and improves the stability of sales forecast accuracy. 3. In the data preprocessing stage, the application designs a missing value filling strategy for continuous and discrete data differentiation, selects an outlier identification and correction method combined with data distribution characteristics, and optimizes data acquisition and storage links through interface protocol adaptation, edge data preprocessing and classified storage, which can improve the refinement degree of raw data processing and provide a more reliable data basis for subsequent prediction; 4. The application divides the data set according to the time sequence characteristics through the model training optimization unit, iteratively updates the sub-model parameters using the gradient descent optimizer, and dynamically adjusts for overfitting, training stagnation and other states, which can reduce the influence of local data distribution deviation on model training, and improve the generalization ability and running stability of the trained model; 5. The application generates sales prediction value and corresponding confidence interval through the prediction result output unit, combines visual presentation, structured export and historical backtracking comparison function, which can intuitively display the prediction result and core feature influence weight, and at the same time, it is convenient for users to check the deviation of prediction result and historical data, which improves the practicality of prediction result and provides effective support for automobile enterprise production scheduling, inventory control and other decisions. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The figure is a schematic diagram of the system framework of the application; Figure 2 The figure is a schematic diagram of the method steps of the application; The meanings of the various reference numerals in the figure are as follows: 100, data acquisition unit; 110, interface protocol adaptation module; 120, edge data preprocessing module; 130, classified storage module; 140, acquisition scheduling module; 150, integrity verification module; 200, data preprocessing unit; 210, missing value filling module; 220, outlier identification and correction module; 230, data conversion module; 240, preprocessing verification module; 300, dynamic feature engineering unit; 310, multi-dimensional feature generation module; 311, time feature extraction submodule; 312, product feature extraction submodule; 313, user feature extraction submodule; 314, environment feature extraction submodule; 320, dynamic screening mechanism module; 321, feature importance quantification submodule; 322, real-time contribution feedback submodule; 323, window iteration update submodule; 324, feature set stability verification submodule; 400, multi-model collaborative prediction unit; 410, time series feature processing module; 411, recurrent neural network submodule; 412, attention mechanism submodule; 420, non-time series feature processing module; 421, tree ensemble learning submodule; 422, feature interaction submodule; 430, model fusion module; 431, adaptive weight adjustment submodule; 432, prediction deviation monitoring submodule; 433, fusion result verification submodule; 500, model training optimization unit; 510, data set time series partitioning module; 520, sub-model parameter iteration module; 530, training termination and adjustment module; 540, training effect verification module; 600, prediction result output unit; 610, prediction result generation module; 620, visualization presentation module; 630, structured export module; 640, historical backtracking comparison module. DETAILED DESCRIPTION
[0018] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] like Figure 1 As shown, this embodiment provides a system for building a multi-dimensional prediction model for automobile sales, including: The data collection unit 100 is used to obtain multi-dimensional automobile-related raw data. By connecting to the automobile production database, sales terminal recording system, user behavior sensing equipment and government open data interface, the data collection unit 100 collects and stores vehicle production records, historical sales data, product technical parameters, user operation trajectories and environmental monitoring data; In this step, the data acquisition unit 100 includes an interface protocol adaptation module 110, an edge data preprocessing module 120, a classification storage module 130, a collection scheduling module 140, and an integrity verification module 150, wherein: The interface protocol adapter module 110 is used to achieve standardized access to multi-source data. By integrating the MQTT protocol adapter component, HTTP protocol adapter component, HTTPS protocol adapter component, and OPCUA protocol adapter component, it connects to the automobile production database, sales terminal recording system, user behavior sensor equipment, and government data disclosure interface respectively, converting the raw data output by heterogeneous data sources into a unified format that can be recognized by the system; It is understandable that automobile production databases typically use an industrial-grade data interaction method based on the OPCUA protocol. The OPCUA protocol adapter component in the interface protocol adapter module 110 establishes a secure channel with the OPCUA server of the database, subscribes to data nodes such as vehicle production work orders and parts supply records, and converts production records transmitted in binary format (such as chassis number generation time, production line workstation completion status, etc.) into ProductionRecord objects defined within the system (including attributes such as batchNumber (production batch number) and productParams (product technical parameter collection)). Most point-of-sale recording systems provide external data interfaces based on HTTP / HTTPS protocols. The HTTP / HTTPS protocol adapter component sends an authenticated GET request to the system's RESTful API to obtain JSON-formatted data containing information such as regional store number, daily sales volume, and sales percentage of vehicle configurations. The data is then parsed into a SalesRecord object (where timeStamp and regionCode are core fields). User behavior sensing devices (such as vehicle-mounted sensors and dealership in-store interactive terminals) use the MQTT protocol for low-power data reporting. The MQTT protocol adapter component acts as an MQTT client to access the device's message queue, subscribe to the userOperation topic, and receive raw data such as user driving habits (such as average speed and charging frequency) and in-store interaction duration, encapsulating it into a UserBehavior object. Government data disclosure interfaces (such as the air quality data interface of the ecological environment department and the road network traffic interface of the transportation department) provide data through the HTTPS protocol. The HTTPS protocol adapter component initiates an SSL / TLS encryption request to obtain environmental monitoring data in JSON or XML format (such as PM2.5 concentration and temperature), and converts it into an EnvironmentData object, so that heterogeneous data from different sources can be recognized by subsequent system modules in a unified object format.
[0020] The edge data preprocessing module 120 is used to process the high-frequency raw data output by the user behavior sensor device, using the LZ4 compression algorithm to reduce the data transmission bandwidth occupancy and the sliding window filtering algorithm to filter out the high-frequency noise during the collection process; Specifically, user behavior sensing devices (for example, in-vehicle sensors) can generate dozens of raw data items per second, including acceleration, steering angle, etc. The edge data preprocessing module 120 first uses the LZ4 compression algorithm to compress continuous raw data blocks. This compresses a 1024-byte raw binary data block to approximately 300-500 bytes (depending on the data's repetition rate). The compressed data is then transmitted to the system server via the network. Furthermore, before transmission, the high-frequency raw data output by the user behavior sensing device is smoothed to filter out high-frequency sensor noise generated by road bumps and other factors during vehicle operation, retaining valid data that reflects user driving behavior trends. Furthermore, this embodiment can set the sliding window size to 15 data points (because car sales are often collected daily, a 15-day window can cover short-term fluctuations over the past two weeks, effectively smoothing out occasional single-day noise; and a weighted average filter is used within the window, with the weights of the most recent three days' data being 0.3, 0.25, and 0.2, respectively, with older data weighted decreasing. This both filters out noise and highlights the impact of recent sales changes).
[0021] It is understandable that the LZ4 compression algorithm selects level 6 of the high compression mode (LZ4 compression levels are divided into "high-speed mode (levels 1-3, focusing on compression speed)" and "high compression mode (levels 4-16, focusing on compression rate)"; this system targets the high-frequency operating data transmitted back by the on-board terminal. When the edge computing power (such as using an 8-core ARM processor with a main frequency of 2.0GHz) allows, it prioritizes ensuring the compression rate to save transmission bandwidth, so level 6 is selected); the window length of the sliding window filtering algorithm is set to 10 data points (determined through comparative experiments: when the window length is 5, the noise filtering is insufficient, when the length is 15, the data transitions smoothly and the details of sales fluctuations are lost, and when the length is 10, it is optimal between noise filtering and detail retention).
[0022] The classification storage module 130 is used for partitioning and storing data according to data characteristics. Specifically, vehicle production records and product technical parameters are stored in a relational database, with the production batch number as the unique index; historical sales data and user operation traces are stored in a time series database, with "timestamp + region code" as the composite index; environmental monitoring data is stored in a distributed file system, with storage paths divided according to the latitude and longitude grid coordinates of the monitoring area; Specifically, the relational database stores vehicle production records (such as production line data of a batch of pure electric vehicles) and product technical parameters (such as the battery capacity and motor power of the vehicle model), with the production batch number as the unique index; the time series database stores historical sales data (such as the sales volume of a certain vehicle model on a certain date in a specific area) and user operation trajectories (such as the user's in-vehicle function operation records at a specific time), with "timestamp + area code (or user ID)" as a compound index; the distributed file system stores environmental monitoring data (such as PM2.5 data at a specific time in a specific area), divides the storage path according to the "monitoring area latitude and longitude grid coordinates / year / month / date / indicator type", and uploads the environmental data files to the corresponding path to achieve distributed storage and rapid retrieval of massive environmental data.
[0023] The collection scheduling module 140 is used to dynamically configure the collection strategy: static data is fully synchronized at a fixed time every day; dynamic data is incrementally collected by triggering the change log; environmental monitoring data is set to a scheduled polling cycle according to the indicator type (every hour for meteorological data and every 15 minutes for traffic flow data); Specifically, the dynamic configuration collection strategy includes: Static data (such as the technical parameters of products corresponding to vehicle production batches, which generally remain stable within the batch production cycle): The collection and scheduling module 140 triggers a full synchronization task at 2:00 a.m. every day through the Linux Cron scheduled task. It calls the relevant components of the interface protocol adaptation module 110 to pull the static data of all production batches as of the previous day and update it to the relational database of the classification storage module 130; Dynamic data (such as real-time sales data from sales terminals, which is continuously updated as transactions are completed): Based on the database change log (such as MySQL's binlog), when a new sales record is generated in the business database of the sales terminal recording system, an incremental collection task is triggered. The collection scheduling module 140 calls a component to read the new or changed entries in the binlog, extract the sales data, and push it to the time series database; Environmental monitoring data: Meteorological data (such as temperature and air pressure) changes relatively slowly. The collection and scheduling module 140 is set to trigger a scheduled poll every hour, and the HTTPS protocol adapter component of the government information disclosure interface is called to obtain data; traffic flow data (such as traffic flow on urban main roads) changes frequently, and a scheduled poll is set to be triggered every 15 minutes to ensure timely acquisition of dynamic traffic environment information.
[0024] The integrity check module 150 is used to verify the integrity of data transmission and storage. It performs transmission verification by calculating the SHA-256 hash value of the data, and performs storage verification by checking the non-empty key fields and the consistency of the number of data entries. A retransmission mechanism is triggered when the verification fails.
[0025] Specifically, verifying the integrity of data transmission and storage includes: Data transmission phase: Before the interface protocol adapter module 110 transmits data to the system, it calculates the SHA-256 hash value for each data object (e.g., ProductionRecord). After serializing the object into a byte array, it generates a hash digest using MessageDigest.getInstance("SHA-256"). The receiving end performs the same operation on the received data. If the hash values are inconsistent, it is determined that the transmission is abnormal, triggering a retransmission mechanism: a retransmission request is sent to the data sender to retrieve the data segment. Data storage stage: After the data is written into the classified storage module 130, the integrity check module 150 performs a non-empty check on the key fields of the relational database, verifies the number of data entries on the time series database (such as confirming whether the number of meteorological data entries within a certain hour matches the number of polling times based on the collection scheduling strategy), and verifies the existence and size of files on the distributed file system; if any of the verifications fails, a data correction process is triggered for the relational database, and a rewrite process is triggered for the time series and file systems to ensure the integrity of the stored data.
[0026] The data preprocessing unit 200 is used to clean and standardize the raw data, fill missing values using statistical interpolation, identify and correct outliers based on data distribution characteristics, and convert unstructured data into standardized data in a unified format using a numerical conversion algorithm; In this step, the data preprocessing unit 200 includes a missing value filling module 210, an outlier identification and correction module 220, a data conversion module 230, and a preprocessing verification module 240, wherein: The missing value filling module 210 is used to process missing items in the original data. Specifically, it uses linear interpolation for continuous numerical data and calculates the filling value based on the trend of valid data points before and after the missing location; and uses the weighted majority method for discrete categorical data, using the category with the highest frequency in the adjacent similar data of the sample where the missing item is located as the filling value, with the frequency weight of recent data being higher than that of remote data. Specifically, the original data is prone to missing items during transmission and storage. The missing value filling module 210 first identifies and classifies the missing items in the original data set, distinguishes between continuous numerical data and discrete categorical data, and then uses corresponding strategies to fill the missing items: For continuous numerical data, the missing value filling module 210 will first locate the missing position, then extract the valid data points before and after the missing position, and calculate the filling value of the missing position based on the changing trend of these data points using linear interpolation, so that the sequence integrity and trend consistency of the continuous data can be maintained.
[0027] For discrete categorical data, the missing value filling module 210 will first determine the range of adjacent similar data of the sample where the missing item is located, count the frequency of occurrence of each category within the range, and assign a higher weight to the recent data. Then, the weighted majority method is used to select the category with the highest frequency of occurrence as the filling value of the missing item, so that the distribution of discrete categorical data is more in line with the category tendency in the actual scenario.
[0028] The outlier identification and correction module 220 is used to detect and correct values that deviate from the data distribution characteristics. Specifically, it includes: for numerical data that follows a normal distribution, outliers are identified by determining whether the deviation of the data point from the mean exceeds 3 times the standard deviation; for non-normally distributed data, the interquartile range method is used; identified outliers are corrected using a local weighted regression method, which corrects the value based on the trend fitting of multiple valid data points before and after the outlier; Specifically, outliers can interfere with the accuracy of subsequent predictions. The outlier identification and correction module 220 first analyzes the distribution characteristics of the data to determine whether the data follows a normal distribution or a non-normal distribution, and then specifically identifies and corrects outliers: If the numerical data follows a normal distribution, the outlier identification and correction module 220 will calculate the mean and standard deviation of the data and identify outliers by determining whether the deviation of the data point from the mean exceeds 3 times the standard deviation; if the data is non-normally distributed, the interquartile range method is used (calculating the interquartile range and identifying outliers based on the range of the range).
[0029] After identifying the outlier, the outlier identification and correction module 220 uses a local weighted regression method to select multiple valid data points before and after the outlier, and fits a correction value that conforms to the overall distribution based on the trend of these data points to replace the original outlier and ensure the rationality of the data distribution.
[0030] The data conversion module 230 is used to achieve data standardization and structural conversion, specifically including: using min-max normalization for numerical data to linearly map the data to a preset interval; using one-hot encoding for categorical data to convert each category into a binary feature vector; using the bag-of-words model to extract keyword features from unstructured data and convert it into a structured frequency matrix; Specifically, in order to enable different types of data to be uniformly processed by subsequent modules, the data conversion module 230 performs standardization and structural conversion on different types of data: For numerical data, the min-max normalization method is used to linearly map the data to a preset numerical interval (such as the [0,1] interval), eliminating the differences between different numerical data due to different dimensions and making the data comparable.
[0031] For categorized data, the one-hot encoding technique is used to convert each category into a binary feature vector containing only 0 and 1, so that the classification information can be recognized by subsequent modules in numerical form.
[0032] For unstructured data (such as text data), the bag-of-words model is used to extract keyword features, count the frequency of keyword occurrence, and convert the unstructured data into a structured frequency matrix to achieve the conversion of unstructured data into numerical features.
[0033] The preprocessing verification module 240 is used to verify the validity of the processed data, specifically including: checking whether the missing values are completely filled, whether the outliers are corrected in accordance with the overall distribution trend of the data, and whether the standardized data is within the preset range. If any condition is not met, the secondary processing of the corresponding module is triggered until all data meet the preprocessing requirements.
[0034] Specifically, to ensure that the processed data meets the requirements of subsequent processes, the pre-processing verification module 240 verifies the validity of the data from multiple aspects: Check whether the missing values have been completely filled. If missing values still exist, trigger the missing value filling module 210 to perform the filling operation again.
[0035] Verify whether the data after outlier correction conforms to the overall distribution trend of the data. If the corrected data still deviates from the overall distribution, trigger the outlier identification and correction module 220 to reprocess.
[0036] Confirm whether the standardized data is within the preset numerical range (for example, whether the value after min-max normalization is within [0,1]), whether the one-hot encoding complies with the rule of "one bit is 1 and the rest are 0", whether the frequency matrix generated by the bag-of-words model is a non-negative integer matrix, etc.; if these conditions are not met, trigger the data conversion module 230 to re-convert the data until all data meet the preprocessing requirements.
[0037] Dynamic feature engineering unit 300 is used to generate a multi-dimensional prediction feature set and adaptively extract key features through a dynamic screening mechanism that quantifies feature importance and provides real-time contribution feedback. The dynamic screening mechanism is based on the calculation of the influence of features on prediction error and combines a sliding window to iteratively update the feature set. In this step, the dynamic feature engineering unit 300 includes a multi-dimensional feature generation module 310, which includes a time feature extraction submodule 311, a product feature extraction submodule 312, a user feature extraction submodule 313, and an environment feature extraction submodule 314, wherein: The time feature extraction submodule 311 is used to extract prediction features from time series data and extract periodic features through the time series period decomposition algorithm. , based on the quantification of the repeated fluctuation pattern of data within a fixed period; extracting trend features through sliding window linear fitting method , based on the change slope of the data in the window to reflect the long-term change direction; the mutation characteristics are extracted by the adjacent window difference threshold method , identify significant time nodes that deviate from the trend; Specifically, the time feature extraction submodule 311 of this embodiment specifically includes: taking "daily historical sales data" as input (recorded as data set , For the Daily sales), extracting periodic features using the STL (Seasonal and Trend decomposition using Loess) time series period decomposition algorithm (which is a conventional algorithm well known to those in the field) , and its calculation formula is ,in For the The daily periodic characteristic value, is the average of all daily sales within a fixed monthly period (i.e. ), the calculation logic is to count the daily sales in the past 12 months, and use this formula to obtain the fluctuation ratio of each day's sales relative to the monthly average to quantify the periodicity; Furthermore, the trend features are extracted by sliding window linear fitting method. , and its calculation formula is ,in is the trend characteristic value (linear regression slope) within the sliding window, is the number of sliding window data points (in this embodiment, ), For the window The calculation logic is to substitute the 7-day sales data and the date number into the formula, and use the positive or negative slope and absolute value to reflect the long-term change direction and significance of sales; Extract mutation features by adjacent window difference threshold method , first calculate the mean difference in sales between two adjacent sliding windows ,in is the average sales volume of the previous window, is the mean sales volume of the next window; the mutation is determined by the following formula: ,in The difference threshold is 1.2 times the maximum value of the mean difference of all adjacent windows in the past three months. The calculation logic is to identify significant time nodes when sales deviate from the trend by comparing the mean difference with the threshold.
[0038] The product feature extraction submodule 312 is used to extract prediction features from technical parameters and extract configuration difference features by calculating the Euclidean distance of parameter vectors. , quantify the differences in hardware parameters of different models; extract performance matching features through parameter combination synergy analysis ,The association relationship based on key parameters reflects the ,adaptability of the product to the usage scenario; Specifically, the product feature extraction submodule 312 takes the “model technical parameter data” as input (recorded as parameter vector) and extracts the configuration difference features by calculating the Euclidean distance of the parameter vector. , and its calculation formula is ,in For car models and car models The configuration difference characteristic value of is the number of technical parameter dimensions (this embodiment includes three parameters: battery capacity, motor power, and wheelbase, namely ), For car models No. Technical parameter values, For car models No. The calculation logic is to substitute the core parameters of the target model and the same-level competitor into the formula, and quantify the difference in hardware parameters between the two through the Euclidean distance. Extract performance matching features through parameter combination synergy analysis , and its calculation formula is ,in For the The performance matching characteristic value of the parameter combination is is the sales volume of the model using this parameter combination, is the total sales volume of all models, is the number of users who need this parameter combination, It is the total number of all users. The calculation logic is to count the sales volume share of the statistical parameter combination and the user share of the corresponding scenario. The ratio of the two reflects the degree of adaptation of the parameter combination to the user demand scenario.
[0039] The user feature extraction submodule 313 is used to extract prediction features from the operation trajectory and extract preference stability features by calculating the entropy value of the operation sequence. , quantify the consistency of users' long-term behavioral habits; extract behavioral conversion characteristics through key path conversion rate analysis , based on the user's operation chain from browsing to ordering, it reflects the stage characteristics of the decision-making process; Specifically, the user feature extraction submodule 313 in this embodiment specifically includes: taking "user operation trajectory data" as input (recorded as operation sequence , For the Operations), extracting preference stability features by calculating the entropy of the operation sequence , and its calculation formula is ,in, is the number of operation categories (this embodiment includes 5 types of operations: browsing models, viewing configurations, consulting customer service, test drive reservations, and placing orders, i.e. ), For the The probability of class operation, and , is the number of such operations, is the total number of operations; the calculation logic is to quantify the consistency of users' long-term behavioral habits through the operation probability and information entropy formula. The lower the entropy value, the more stable the preference. Extract behavioral conversion characteristics through key path conversion rate analysis , and its calculation formula is ,in Transform the characteristic value of the user's key path behavior, is the number of key path nodes (the path in this embodiment is "browse-view configuration-consult-test drive-order", i.e. , To complete the The number of users operating on each node, To complete the The number of users operating on each node, Represents a multiplication symbol; the calculation logic is to sequentially calculate the conversion rates of adjacent nodes and multiply them together, reflecting the stage-by-stage advancement efficiency of the user's purchase decision-making process.
[0040] The environmental feature extraction submodule 314 is used to extract prediction features from environmental data and generate regional correlation features by mining the co-occurrence patterns of meteorological, traffic data and historical sales. , quantify the impact of environmental factors on sales in different regions.
[0041] Specifically, the environmental feature extraction submodule 314 includes: taking "regional environmental data + regional sales data" as input, generating regional correlation features by mining the co-occurrence patterns of meteorological, traffic data and historical sales , and its calculation formula is ,in For a specific environment The regional correlation eigenvalue under For the regional environment Sales volume under rainy conditions (such as sales volume under rainy conditions), is the average monthly sales volume for the region, and , For the region Daily sales; the calculation logic is to statistically calculate the ratio of regional sales under a specific environment to the monthly average sales, and quantify the impact of environmental factors on sales in different regions through the size of the ratio. A ratio greater than 1 indicates that the environment has a promoting effect on sales, while a ratio less than 1 indicates an inhibitory effect.
[0042] It is understandable that regional correlation characteristics This data can be obtained through the following methods: connecting to regional meteorological platforms (such as China Weather Network API) and real-time data interfaces of transportation departments (such as AutoNavi traffic congestion data), collecting historical meteorological data (temperature, precipitation, etc.), traffic data (congestion index, traffic volume) and daily automobile sales data of the corresponding region for the past three years; using association rule mining algorithms (such as Apriori algorithm) to mine the co-occurrence patterns of meteorological, traffic and historical sales such as "high temperature + evening peak congestion - increased sales of family cars", and calculating the co-occurrence frequency and confidence as the basis for the prediction. Numerical representation of .
[0043] In this step, the dynamic feature engineering unit 300 further includes a dynamic screening mechanism module 320, which includes a feature importance quantification submodule 321, a real-time contribution feedback submodule 322, a window iterative update submodule 323, and a feature set stability verification submodule 324, wherein: The feature importance quantification submodule 321 is used to calculate the influence of the prediction features extracted by the multi-dimensional feature generation module 310 on the prediction results. ,It is determined by comparing the difference in prediction error after including the prediction feature and randomly replacing the prediction feature value.,The influence value is positively correlated with the error change amplitude; Specifically, the feature importance quantification submodule 321 takes the initial feature set extracted by the multi-dimensional feature generation module 310 and the historical sales labels as input to calculate the influence of each prediction feature on the prediction result. , the calculation logic is as follows: First, the initial feature set and sales volume labels are input into the basic prediction model (such as linear regression model), and the prediction error of the model is calculated (using the mean absolute error MAE) and recorded as ; Then, keeping other features unchanged, randomly disrupt the numerical order of the target feature to destroy its correlation with sales, input the basic prediction model again and calculate the new prediction error ; Finally Determine the influence of the feature, where is the characteristic influence value, which is positively correlated with the error change amplitude. is the prediction error of the original feature set, In order to disrupt the prediction error after the target feature, this logic can be used to quantify the influence of each feature on the prediction result. The greater the influence, the higher the importance of the feature.
[0044] The real-time contribution feedback submodule 322 is used to monitor the actual role of various prediction features in recent predictions. , based on the weighted cumulative value of feature importance within a sliding window (the window length is dynamically adjusted according to the prediction period). The weight coefficient of recent data in the window is higher than that of long-term data to ensure the timeliness of feedback; Specifically, the dynamic screening mechanism module 320 of this embodiment is based on the feature influence output by the feature importance quantification submodule 321 to monitor the actual role of various prediction features in recent predictions. , the calculation logic is: First, dynamically adjust the sliding window length according to the forecast period (for example, if the forecast period is weekly, the window length is set to 4 weeks); Then, data from different time periods within the window are assigned decreasing weights (the weight of the recent week is set to 0.4, the previous week is set to 0.3, the week before that is set to 0.2, and the fourth week is set to 0.1) to ensure that recent data plays a dominant role in the contribution; Finally, the real-time contribution is calculated based on the weighted cumulative value of the feature importance in the sliding window, that is, ,in is the real-time contribution of the feature, is the feature influence in the past week, is the feature influence of the previous week, is the feature influence of the previous week, This is the feature influence in the fourth week at the farthest point in time. This logic can reflect the actual effect of the feature in the recent forecast in real time.
[0045] The window iteration update submodule 323 is used to dynamically adjust the feature set. When the influence of any prediction feature extracted by the multi-dimensional feature generation module 310 is Two consecutive windows are below the preset threshold, or their contribution When three consecutive windows show a monotonically decreasing trend, the prediction feature is automatically removed from the current feature set; at the same time, a new prediction feature is generated by combining cross-dimensional features. When the initial influence of the new prediction feature is When it exceeds the average level of the current feature set, it is included in the feature set to form an updated feature set; Specifically, the window iteration update submodule 323 in this embodiment uses the influence of the feature importance quantification submodule 321 The contribution of the real-time contribution feedback submodule 322 Based on this, it is used to dynamically adjust the feature set. The specific logic includes two parts: removing inefficient features and incorporating new features. The logic of eliminating inefficient features is as follows: First, based on the historical feature influence data, a preset threshold is set (50% of the average influence of all features in the past three months). Two consecutive sliding windows are below the preset threshold, or their contribution If three consecutive sliding windows show a monotonically decreasing trend, the feature is automatically removed from the current feature set; the logic for incorporating new features is to generate new prediction features by combining cross-dimensional features (such as and Combined into "region-period joint features", and user characteristics Combined into "user preference-product configuration matching feature"), calculate the initial impact of the new feature ,If the initial influence exceeds the average influence level of the current feature set, it will be included in the feature set, and finally form the updated feature set.
[0046] The feature set stability check submodule 324 is used to verify the validity of the updated feature set by calculating the overlap rate between the current feature set before the update and the feature set generated after the update by the calculation window iteration update submodule 323. and the forecast error fluctuation value To evaluate, when the overlap ratio Lower than the preset ratio or error fluctuation value When the threshold is exceeded, the multi-dimensional feature generation module 310 is triggered to re-extract prediction features to optimize the feature set.
[0047] Specifically, the feature set stability verification submodule 324 of this embodiment uses the updated feature set output by the window iterative update submodule 323 as input to verify the validity of the updated feature set. The specific logic includes three parts: overlap rate calculation, prediction error fluctuation value calculation, and validity determination. The overlap rate calculation logic is to count the number of common features between the feature set before and after the update, and calculate the overlap rate using the following formula: ,in is the feature set overlap rate, is the number of common features, is the number of features before updating, is the number of features after update; The prediction error fluctuation value calculation logic is to input the feature sets before and after the update into the same prediction model, calculate the error (MAE) of the two predictions, and calculate the error fluctuation value using the following formula: ,in is the forecast error fluctuation value, is the prediction error of the updated feature set, is the prediction error of the feature set before updating; The validity judgment logic is to set the preset standard (feature set overlap rate Not less than 60%, forecast error fluctuation value Less than or equal to 10%), if the updated feature set meets the standard, the verification is determined to be passed and the feature set is output; if it does not meet the standard (such as the overlap rate is less than 60% or the error fluctuation value exceeds 10%), the multi-dimensional feature generation module 310 is triggered to re-extract the prediction features (such as adjusting the time window length, supplementing the product parameter dimension) until the generated feature set passes the stability check.
[0048] Multi-model collaborative prediction unit 400 is used to build a time series and non-time series dual-channel prediction framework. It uses a recurrent neural network with an attention mechanism to process time series features, a tree-structured ensemble learning algorithm to process product attributes and user behavior characteristics, and an adaptive weight adjustment algorithm based on real-time prediction deviations to achieve multi-sub-model output fusion. In this step, the multi-model collaborative prediction unit 400 includes a time series feature processing module 410 and a non-time series feature processing module 420, wherein: The time series feature processing module 410 is used to process time series features, and includes a recurrent neural network submodule 411 and an attention mechanism submodule 412, wherein: The recurrent neural network submodule 411 adopts a long short-term memory network structure and uses the periodic features generated by the time feature extraction submodule 311 to extract the periodic features. , trend characteristics and mutation characteristics It takes the input as input, captures the dependencies of different time scales through the gating mechanism of input gate, forget gate and output gate, and outputs the hidden layer representation sequence of temporal features; Specifically, the recurrent neural network submodule 411 of this embodiment specifically includes: using a long short-term memory network structure to 、 、 is the input (input feature vector , The forget gate controls the proportion of historical information retained, the input gate generates candidate cell states (mapped to [-1, 1] by tanh activation) and incorporates current information, and the output gate controls the output proportion of the cell state; The core calculation logic is: implemented through the cell state update formula and the hidden state formula. The cell state update formula is: , the hidden state formula is: ; wherein is the time step cell state, is an element-wise multiplication, is the history cell state, represents the candidate cell state at the time step, is the time step hidden state, is the output of the forget gate, is the output of the input gate, represents the output of the output gate, represents the hyperbolic tangent activation function; and the final generated hidden layer representation sequence ( is the total number of time steps), is output to the attention mechanism submodule 412; wherein the gate calculation is based on a "weight matrix + sigmoid activation function" (the dimension of the weight matrix is "the number of hidden layer neurons x (the number of hidden layer neurons + 3)", and 3 is the input feature dimension), which ensures that the gate output value range is [0, 1].
[0049] The attention mechanism submodule 412 is used to enhance the influence weight of the key timing feature, and the attention weight of each time step is calculated based on the hidden layer representation sequence , which gives higher weight to the time step containing the mutation feature and the recent time step within the prediction window, and converts the hidden layer representation sequence into an aggregated vector of timing features through weighted summation as the intermediate prediction result of the timing channel; Specifically, the attention mechanism submodule 412 in the embodiment is used to enhance the influence weight of the key timing feature, taking the hidden layer representation sequence as input, first calculating the importance score of each time step through additive attention (based on the global weight vector, the attention weight matrix, and the hidden layer representation of the starting time step of the prediction window), and then obtaining the weight of each time step through the attention weight normalization formula, which is expressed as: ; in the formula is the attention weight of the time step, is the basic attention score, , is the weight adjustment coefficient (value [0.5, 1.5], which can be set to , This design can make the attention mechanism better capture the key information in the timing of automobile sales: It can strengthen the attention to the mutation of sales (such as the change point of promotion and policy impact), It can properly emphasize the importance of the recent time step (recent sales). The combination of the two allows the model to more accurately extract effective features and improve sales forecast accuracy). Marks the recent time step (1 for the first 5 time steps of the prediction window, 0 otherwise), represents the natural exponential function, is the total number of time steps, Represents the time step index (traverses all time steps to calculate the denominator of weight normalization); finally, the weighted sum Generate time series feature aggregation vector (i.e., the intermediate prediction results of the time series channel), strengthening and the information of the recent time step are output to the model fusion module 430.
[0050] The non-temporal feature processing module 420 is used to process product attributes and user behavior features, and includes a tree-structured ensemble learning submodule 421 and a feature interaction submodule 422, wherein: The tree ensemble learning submodule 421 uses the gradient boosting tree algorithm to extract the configuration difference features generated by the product feature extraction submodule 312. , performance matching characteristics , and the preference stable features generated by the user feature extraction submodule 313 , behavioral conversion characteristics As input, a nonlinear prediction model is constructed through iterative training of multiple decision trees, and the basic prediction value of non-time series features is output; Specifically, the tree ensemble learning submodule 421 adopts the gradient boosting tree algorithm to 、 、 、 is the input (input feature vector ), initial model Set to sample mean (squared loss ), is the total number of historical sales samples in the "Region-Date" dimension. For the actual sales volume of the sample); In each iteration, the prediction residual of the previous model is calculated first, and then the residual is fitted using a decision tree (the tree depth is set to 5 and the number of leaf nodes is set to 10). The contribution of a single tree is controlled by the integrated model update formula, which is: ,in For the front Tree ensemble model, is the learning rate. In this embodiment, , For the A decision tree, Before An ensemble model of 100 decision trees (prediction results of the previous iteration); after 100 iterations, the basic prediction value formula is used to output the basic prediction value of non-time series features, capturing the nonlinear correlation between product and user characteristics and sales. The expression is: ,in Represents the basic prediction value of non-time series features (the prediction result when high-order interaction features are not added).
[0051] The feature interaction submodule 422 is used to mine cross-dimensional feature associations, using the regional association features generated by the environmental feature extraction submodule 314 As a link, calculate the interaction strength between product features and user features , generate high-order interaction features, and input the high-order interaction features into the tree-structured ensemble learning submodule 421 for secondary training, and correct the basic prediction value to form the final output of the non-time series channel.
[0052] Specifically, the feature interaction submodule 422 is To mine cross-dimensional feature associations for ties, first calculate product features using the interaction strength formula (This embodiment takes ) and user characteristics (This embodiment takes ) interaction strength , the specific expression is: ; Where, is the probability distribution (estimated by historical sample frequency, such as ), the larger the value, the more significant the cross-dimensional correlation; Represents the characteristics of a given environment When the product features and user characteristics The interaction intensity; They are 、 The specific value of represents the natural logarithm function; Then, product and user characteristics are respectively Multiply to generate high-order interactive features, add the original feature set to retrain GBDT, and obtain the final output of the non-sequential channel , capturing the cross-dimensional association of “product-user-environment” and outputting it to the model fusion module 430 .
[0053] In this step, the multi-model collaborative prediction unit 400 further includes a model fusion module 430, which includes an adaptive weight adjustment submodule 431, a prediction deviation monitoring submodule 432, and a fusion result verification submodule 433, wherein: The adaptive weight adjustment submodule 431 is used to dynamically fuse the output results of the time series feature processing module 410 and the non-time series feature processing module 420, and calculate the weight coefficient based on the real-time prediction deviation. : Time series model weight and non-series model weights The sum is 1, the weight value is generated by the inverse deviation function, and the weight update cycle is consistent with the prediction cycle; Specifically, the adaptive weight adjustment submodule 431 is used to dynamically fuse the time series and non-time series channel outputs, and the time series channel intermediate prediction results Mapped to sales forecast value through the fully connected layer , non-sequential channel input is Based on the real-time deviation (timing deviation) provided by the prediction deviation monitoring submodule 432 , non-timing deviation , are the square differences between the predicted value and the actual sales volume), through the weight formula , and the weight formula , calculate the weight coefficient; where, , the smaller the deviation, the greater the weight; represents the weight coefficient of the time series model, Represents the weight coefficient of the non-time series model; the final prediction value is generated through the fusion formula , whose expression is: , the weight update cycle is consistent with the prediction cycle to ensure adaptation to the latest prediction error.
[0054] The prediction deviation monitoring submodule 432 is used to quantify the submodule output error and calculate the time series deviation of the intermediate prediction result of the time series feature processing module 410 based on the actual sales data. (square difference between the predicted value and the actual value) and the non-time series deviation finally output by the non-time series feature processing module 420 , the deviation sequence is stored through a sliding window with a length of 5 prediction periods to provide a historical basis for weight adjustment; Specifically, the forecast deviation monitoring submodule 432 is used to quantify the submodule output error and store historical deviations. After each forecast cycle, the actual sales volume is used to monitor the output error. Taking the prediction deviation of the time series model and the non-time series model as the benchmark, the prediction deviation (both are the squared difference between the predicted value and the actual value) of the time series model and the non-time series model is calculated; the deviation is stored in a sliding window with a length of 5 prediction periods to form a historical deviation sequence (the earliest period deviation is eliminated when the new period deviation is added), providing a multi-period error reference for adaptive weight adjustment and avoiding weight fluctuations caused by single-period abnormal deviations.
[0055] The fusion result verification submodule 433 is used to verify the reliability of the fusion prediction by calculating the mean absolute error between the fusion result and the actual sales volume. And the improvement rate of the fusion result compared with the output of the time series feature processing module 410 and the non-time series feature processing module 420 Assessment: When Exceeds the preset threshold or When , the adaptive weight adjustment submodule 431 is triggered to recalculate the weights, and the deviation sensitive features are fed back to the time series feature processing module 410 and the non-time series feature processing module 420 to drive the secondary optimization of the feature weights.
[0056] Specifically, the fusion result verification submodule 433 is used to verify the reliability of the fusion prediction, using the fusion results of 5 prediction cycles. , actual sales The time series and non-time series prediction values are input, and the fusion error is quantified by the mean absolute error formula ( The smaller the prediction, the more accurate it is). The formula is: The fusion advantage is evaluated by the improvement rate formula, which is: ;in, , 、 are the mean absolute errors of the time series and non-time series models respectively; Represents the minimum mean absolute error of a single model; Indicates the number of prediction cycles in the verification window (take the prediction data of the most recent 5 cycles); Indicates the forecast period index (traversing 5 periods); Indicates the The fusion prediction value of the cycle; Indicates the Actual sales volume for each cycle; represents the absolute value function (eliminating the positive and negative cancellation of errors); Indicates the mean absolute error of the fusion result; like ( Based on historical optimal preset threshold) or , the adaptive weight adjustment submodule 431 is triggered to recalculate the weights and feedback the deviation-sensitive features to the time series feature processing module 410 and the non-time series feature processing module 420 to drive the secondary optimization of the feature weights and ensure the reliability of the fusion results.
[0057] The model training optimization unit 500 is used to perform parameter training and performance optimization on the sub-models of the multi-model collaborative prediction unit 400. The training set, validation set, and test set are divided based on the historical timestamps corresponding to the prediction features output by the dynamic feature engineering unit 300. The core parameters of the sub-models are iteratively updated using a gradient descent optimizer. The training is terminated when the validation set error does not improve for consecutive preset rounds. If overfitting occurs or the error does not improve, the sub-model parameters are adjusted or some data is supplemented before restarting the training. In this step, the model training optimization unit 500 includes a data set time series partitioning module 510, a sub-model parameter iteration module 520, a training termination and adjustment module 530, and a training effect verification module 540, wherein: The dataset time series partitioning module 510 is used to partition the training data according to the time series characteristics. Based on the historical timestamps corresponding to the prediction features output by the dynamic feature engineering unit 300, the dataset is divided into a training set, a validation set, and a test set in chronological order. When partitioning, ensure that the training set covers at least 2 complete The corresponding historical data, validation set, and test set each cover at least 1 complete Corresponding historical data, while avoiding mutation characteristics Concentrated time intervals to avoid data distribution shifts that affect training stability; Specifically, the dataset time series partitioning module 510 divides the dataset into a training set, a validation set, and a test set in chronological order based on the historical timestamps corresponding to the prediction features output by the dynamic feature engineering unit 300. The partitioning process must meet three core requirements: First, data coverage integrity, ensuring that the training set covers at least 2 complete The corresponding historical data, validation set, and test set each cover at least 1 complete Corresponding historical data to ensure that the model learns complete temporal patterns; The second is data distribution stability, strictly avoiding mutation characteristics A concentrated time interval to avoid data distribution deviation within this interval affecting the stability of the trained model; The third is the uniqueness of the division logic. The time sequence is used as the only basis for division throughout the process, and the time sequence is not disrupted to ensure that the training set, verification set, and test set are consistent with the time logic in the actual business scenario.
[0058] The sub-model parameter iteration module 520 is used to perform the sub-model core parameter update, and uses the gradient descent optimizer to optimize the two types of sub-models of the multi-model collaborative prediction unit 400: for the recurrent neural network with attention mechanism of the time series feature processing module 410, optimize its network connection weights, attention weights, and the like. Calculate coefficients and recurrent layer gating biases; optimize the decision tree feature splitting threshold and node splitting gain coefficient for the tree-structured ensemble learning model of the non-time-series feature processing module 420; dynamically set the initial learning rate based on the complexity of the sub-model, decay at a fixed rate for each preset training round, and simultaneously record the error changes of the validation set during the decay process; Specifically, the sub-model parameter iteration module 520 specifically includes: using a gradient descent optimizer to update the core parameters of the two types of sub-models of the multi-model collaborative prediction unit 400 respectively to ensure that the parameters are adapted to the prediction requirements; For the attention mechanism recurrent neural network of the time series feature processing module 410, the following three core parameters are optimized: network connection weight (controls the strength of information transmission between neurons), attention weight Calculate coefficients and loop layer gating bias (adjust the ratio of information forgotten and retained in the loop layer); For the tree-based ensemble learning model of the non-time-series feature processing module 420, the following two core parameters are optimized: the decision tree feature splitting threshold (the critical value for determining whether a feature should be split) and the node splitting gain coefficient (an indicator for measuring the information gain after feature splitting); During the parameter iteration process, the initial learning rate is dynamically set based on the complexity of the sub-model (for example, if the complexity of the time series sub-model is high, the initial learning rate is set to a smaller value). After each preset training round (for example, 10 rounds), the learning rate is reduced at a fixed ratio (for example, 10% decay each time). During the decay process, the changes in the validation set error are synchronously recorded to provide a basis for subsequent training adjustments.
[0059] It is understandable that using the Adam optimizer as a gradient descent optimizer, compared to SGD (stochastic gradient descent), Adam can adaptively adjust the learning rate based on the first-order moment estimation and second-order moment estimation of the gradient. It is more suitable for the parameter update scenario of the "temporal LSTM + non-temporal gradient boosting tree" multi-submodel collaborative training, with faster convergence and more stable training process.
[0060] The training termination and adjustment module 530 is used to trigger the termination of training and handle abnormal training states: the preset termination condition is "the decrease in the validation set error for 3-5 consecutive rounds is lower than the preset error threshold", and training is terminated when the condition is met; if the overfitting state of "the training set error continues to decrease while the validation set error increases" occurs, the time series sub-model is adjusted by increasing the dropout probability, and the non-time series sub-model is adjusted by increasing the tree pruning intensity; if the errors of the training set and validation set do not improve for multiple consecutive rounds, less than or equal to 20% of recent data are selected from the test set to supplement the training set, and the parameter iteration process is restarted; Specifically, the training termination and adjustment module 530 is responsible for triggering the training termination process and handling abnormal training status to ensure a balance between training efficiency and model performance. Training termination follows preset conditions, including: when the validation set error decreases below a preset error threshold for 3-5 consecutive rounds (e.g., the decrease is less than 0.5% per round), the model performance is determined to be stable and training is terminated. Two types of abnormal training states are handled separately: one is the overfitting state. When "the error of the training set continues to decrease while the error of the validation set increases", the time series sub-model is adjusted by increasing the Dropout probability (reducing the risk of over-dependence of neurons), and the non-time series sub-model is adjusted by increasing the tree pruning intensity (reducing redundant branches of the decision tree); the other is the double error no improvement state. When the errors of both the training set and the validation set have not improved for multiple consecutive rounds (such as 5 rounds), less than or equal to 20% of recent data are selected from the test set to supplement the training set (recent data is closer to the current business scenario), and the sub-model parameter iteration process is restarted.
[0061] The training effect verification module 540 is used to verify the effectiveness of the trained sub-model. It uses the test set that does not participate in parameter iteration as input and calculates the mean absolute percentage error between the predicted value and the actual sales volume. , goodness of fit ;when and When the preset model performance threshold is met, the training is determined to be effective and the final sub-model parameters are output; if the threshold requirement is not met, the module returns to the data set time series partitioning module 510 to readjust the partitioning ratio and execute the training process again.
[0062] Specifically, the training effect verification module 540 takes the test set that does not participate in parameter iteration as input and verifies the effectiveness of the trained sub-model through quantitative indicators. The core verification indicators include the mean absolute percentage error , goodness of fit ( Used to measure the degree of fit between the forecast value and the actual sales); the verification logic is: when and All meet the preset model performance threshold (such as ), the training is determined to be valid, and the final sub-model parameters are output to the multi-model collaborative prediction unit 400; if any indicator does not meet the threshold requirement, the data set time series partitioning module 510 is returned to readjust the partition ratio of the training set, validation set, and test set (such as appropriately increasing the proportion of the training set), and the complete training process is executed again.
[0063] The prediction result output unit 600 is used to generate and output the automobile sales prediction results. By combining the trained model parameters output by the model training optimization unit 500 and the real-time features of the dynamic feature engineering unit 300, the sales prediction value and confidence interval for the future preset period are generated. The prediction value, confidence interval and core feature influence weight are presented through a visual interface, and the structured export of the prediction results is supported. The historical sales data stored in the data acquisition unit 100 can be called to realize the retrospective comparison between the prediction results and the actual historical sales.
[0064] In this step, the prediction result output unit 600 includes a prediction result generation module 610, a visualization presentation module 620, a structured export module 630, and a historical backtracking comparison module 640, wherein: The prediction result generation module 610 is used to generate sales forecast data, receive the trained model parameters output by the model training optimization unit 500, load the real-time features output by the dynamic feature engineering unit 300, calculate and output the sales forecast value for the corresponding period according to the preset prediction period, and calculate the confidence interval corresponding to the prediction value based on the model prediction error characteristics; Specifically, the prediction result generation module 610 receives the trained model parameters (time series sub-model weight, non-time series sub-model split threshold, etc.) of the model training optimization unit 500, loads the real-time features (real-time 、 、 etc.); calculate the sales forecast value of the corresponding period according to the preset period (day / week / month, such as "week"), and based on the historical forecast error of the model (refer to the error range reflected by MAPE), use " ” Calculate the 95% confidence interval and output the “period-predicted value-confidence interval” data.
[0065] The visualization module 620 is used to display forecast-related information. It simultaneously presents sales forecast values, corresponding confidence intervals, and the impact weights of core forecast features through the interface. The forecast values and confidence intervals are displayed in the form of trend charts, and the impact weights of core features are quantitatively presented in simple charts. Specifically, the interface of the visualization module 620 simultaneously displays the following three types of information: predicted values and confidence intervals (the horizontal axis is the forecast period, the vertical axis is the sales volume, the forecast value is presented with a solid line, and the confidence interval is presented with a shaded trend chart), the core feature influence weight (the top 5 high-contribution features, quantified with a horizontal bar chart); annotation of the forecast period type, confidence level, and model version, and support for hovering the mouse to view specific values.
[0066] The structured export module 630 is used to output the forecast result file. It supports exporting sales forecast values, confidence intervals, and core feature influence weights in a standardized structured format. During the export process, the integrity of key data fields is automatically verified to ensure that there are no missing data. Specifically, the structured export module 630 supports export in Excel / CSV format, including the fields of "forecast period, forecast value, upper and lower limits of confidence interval, and core feature weight"; the integrity of key fields (forecast value, period) is automatically checked during export, and if missing, corrections will be prompted. After the verification is passed, the file will be named according to "Automobile Sales Forecast_Period_Date", and local or enterprise system storage is supported.
[0067] The historical backtracking comparison module 640 is used to compare the forecast with historical data, call the historical sales data stored in the data acquisition unit 100, align the forecast results with the historical actual sales data according to the time dimension corresponding to the forecast period, calculate the deviation value between the two, and present the backtracking results in the form of a table or comparison chart.
[0068] Specifically, the historical retrospective comparison module 640 calls the historical sales data of the data collection unit 100, aligns it according to the forecast period time dimension (such as weekly forecast corresponds to historical weekly data for the same period); calculates the deviation between the forecast value and the actual historical sales (deviation = forecast value - historical value); presents the results in a comparison table (listing period, historical sales, forecast value, deviation) or a dual-axis chart (historical sales bar chart, forecast value line chart), and supports filtering segmented data by region / model.
[0069] like Figure 2 As shown, this embodiment also provides a method for constructing a multi-dimensional prediction model for automobile sales, based on the above-mentioned multi-dimensional prediction model construction system for automobile sales, including the following steps: S100, Multi-source Automotive Raw Data Collection: Connect to automotive production databases, point-of-sale recording systems, user behavior sensing devices, and government data disclosure interfaces through interface protocol adaptation; process high-frequency raw data from user behavior sensing devices through edge data preprocessing; store data in classified storage partitions based on data characteristics; dynamically configure collection strategies through collection scheduling; verify data transmission and storage integrity through integrity checks, with verification failures triggering retransmission; S200, Data Cleaning and Standardization: Missing items in the original data are processed through missing value filling. Continuous numerical data uses linear interpolation, and discrete categorical data uses the weighted majority method. Outliers are detected and corrected through outlier identification and correction. Normally distributed data is identified based on 3 times the standard deviation, and non-normally distributed data is identified using the interquartile range method. Outliers are corrected using local weighted regression. Data standardization and structural transformation are achieved through data conversion. Numerical data uses min-max standardization, categorical data uses one-hot encoding, and unstructured data uses the bag-of-words model. The validity of the processed data is verified through preprocessing verification. S300, Dynamic Feature Set Generation: Extract multi-dimensional prediction features through multi-dimensional feature generation; screen features through a dynamic screening mechanism, calculate feature influence and real-time contribution, and iteratively update the feature set; verify the validity of the updated feature set through feature set stability check, and re-extract features if it does not meet the requirements; S400, Dual-channel Collaborative Prediction and Fusion: Performs temporal feature processing on time series features, employs recurrent neural networks to capture time-scale dependencies, and incorporates attention mechanisms to enhance the impact of key temporal features. It also performs non-temporal feature processing on product attributes and user behavior features, constructs a prediction model using tree-structured ensemble learning, and mines cross-dimensional feature associations through feature interaction. It also fuses the outputs of the two channels, calculates adaptive weights, monitors prediction deviations, verifies the fusion results, and optimizes them. S500, Model Training Optimization: The training set, validation set, and test set are divided into time series by data set time series to ensure that the complete cycle data is covered and the intervals with concentrated mutation features are avoided; the sub-model parameters are updated by gradient descent optimizer through sub-model parameter iteration, and the initial learning rate is dynamically set and decayed by rounds; training termination is triggered by training termination and adjustment to deal with overfitting and training stagnation; the effectiveness of the sub-model after training is verified by training effect verification. If it does not meet the requirements, the data set division is readjusted and training is resumed; S600, forecast result output: Generate sales forecast value and corresponding confidence interval based on forecast results; display forecast value, confidence interval and core feature influence weight through visualization; export forecast results in standardized format through structured export and verify data integrity; align forecast results with historical sales data through historical backtracking comparison, calculate deviation and present comparison results.
[0070] Those skilled in the art will appreciate that the process of implementing all or part of the steps of the above embodiments may be accomplished by hardware, or by instructing related hardware through a program.
[0071] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-dimensional prediction model construction system for automobile sales, characterized by: include: A data acquisition unit (100), the data acquisition unit (100) is used to obtain multi-dimensional automobile-related raw data, and collects and stores vehicle production records, historical sales data, product technical parameters, user operation trajectories and environmental monitoring data by connecting to an automobile production database, a sales terminal recording system, a user behavior sensing device and a government open data interface; A data preprocessing unit (200), the data preprocessing unit (200) is used to clean and standardize the original data, fill missing values using statistical interpolation, identify and correct outliers based on data distribution characteristics, and convert unstructured data into standardized data in a unified format using a numerical conversion algorithm; A dynamic feature engineering unit (300), the dynamic feature engineering unit (300) is used to generate a multi-dimensional prediction feature set, and adaptively extract key features through a dynamic screening mechanism of feature importance quantification and real-time contribution feedback, the dynamic screening mechanism is based on the calculation of the influence of the feature on the prediction error, combined with a sliding window to iteratively update the feature set; A multi-model collaborative prediction unit (400) is used to build a time series-non-time series dual-channel prediction framework, using a recurrent neural network with an attention mechanism to process time series features, processing product attributes and user behavior features through a tree-like ensemble learning algorithm, and realizing multi-sub-model output fusion based on an adaptive weight adjustment algorithm based on real-time prediction deviation; A model training optimization unit (500) is used to perform parameter training and performance optimization on the sub-models of the multi-model collaborative prediction unit (400), by dividing the training set, the validation set, and the test set based on the historical timestamp corresponding to the prediction feature output by the dynamic feature engineering unit (300), and iteratively updating the core parameters of the sub-model using a gradient descent optimizer. When the validation set error does not improve for consecutive preset rounds, the training is terminated. If overfitting occurs or the error does not improve, the sub-model parameters are adjusted or some data is supplemented before restarting the training. A prediction result output unit (600) is used to generate and output automobile sales forecast results, generate sales forecast values and confidence intervals for a future preset period by combining the trained model parameters output by the model training optimization unit (500) with the real-time features of the dynamic feature engineering unit (300), present the forecast values, confidence intervals, and core feature influence weights through a visual interface, support structured export of forecast results, and call historical sales data stored in the data acquisition unit (100) to achieve retrospective comparison between forecast results and historical actual sales.
2. The automobile sales multi-dimensional prediction model construction system according to claim 1 is characterized in that: The data acquisition unit (100) comprises an interface protocol adaptation module (110), an edge data pre-processing module (120), a classification storage module (130), an acquisition scheduling module (140) and an integrity verification module (150), wherein: The interface protocol adaptation module (110) is used to achieve standardized access to multi-source data, and by integrating an MQTT protocol adaptation component, an HTTP protocol adaptation component, an HTTPS protocol adaptation component, and an OPCUA protocol adaptation component, it is connected to the automobile production database, the sales terminal recording system, the user behavior sensor device, and the government affairs disclosure data interface, respectively, to convert the raw data output by the heterogeneous data sources into a unified format that can be recognized by the system; The edge data preprocessing module (120) is used to process the high-frequency raw data output by the user behavior sensing device, using the LZ4 compression algorithm to reduce the data transmission bandwidth occupancy, and filtering out the high-frequency noise in the collection process through the sliding window filtering algorithm; The classification storage module (130) is used for partitioning and storing data according to data characteristics, specifically including: storing vehicle production records and product technical parameters in a relational database, with the production batch number as the unique index; storing historical sales data and user operation traces in a time series database, with "timestamp + region code" as the composite index; storing environmental monitoring data in a distributed file system, with storage paths divided according to the latitude and longitude grid coordinates of the monitoring area; The acquisition scheduling module (140) is used to dynamically configure the acquisition strategy: static data adopts full synchronization at a fixed time every day; dynamic data adopts incremental acquisition triggered by change logs; environmental monitoring data sets a regular polling cycle according to the indicator type; The integrity check module (150) is used to verify the integrity of data transmission and storage, performs transmission verification by calculating the SHA-256 hash value of the data, performs storage verification by checking the key field non-empty state and the consistency of the number of data items, and triggers a retransmission mechanism when the verification fails.
3. The automobile sales multi-dimensional prediction model construction system according to claim 2 is characterized in that: The data pre-processing unit (200) comprises a missing value filling module (210), an abnormal value identification and correction module (220), a data conversion module (230) and a pre-processing verification module (240), wherein: The missing value filling module (210) is used to process missing items in the original data, specifically including: using linear interpolation for continuous numerical data, and calculating the filling value based on the trend of valid data points before and after the missing position; using weighted majority method for discrete categorical data, and using the category with the highest frequency of occurrence in the adjacent similar data of the sample where the missing item is located as the filling value, and the frequency weight of recent data is higher than that of long-term data; The outlier identification and correction module (220) is used to detect and correct values that deviate from data distribution characteristics, specifically including: for numerical data that obeys normal distribution, identifying outliers by judging whether the deviation between the data point and the mean exceeds 3 times the standard deviation; for non-normal distribution data, using the interquartile range method; correcting the identified outliers by using the local weighted regression method, and correcting the values based on the trend fitting of multiple valid data points before and after the outlier; The data conversion module (230) is used to achieve data standardization and structured conversion, specifically including: using min-max standardization for numerical data to linearly map the data to a preset interval; using one-hot encoding for categorical data to convert each category into a binary feature vector; using a bag-of-words model to extract keyword features from unstructured data and convert it into a structured frequency matrix; The pre-processing verification module (240) is used to verify the validity of the processed data, specifically including: checking whether the missing values are completely filled, whether the outliers are corrected in accordance with the overall distribution trend of the data, and whether the standardized data is within a preset range. If any of the conditions is not met, the secondary processing of the corresponding module is triggered until all the data meet the pre-processing requirements.
4. The automobile sales multi-dimensional prediction model construction system according to claim 3 is characterized in that: The dynamic feature engineering unit (300) includes a multi-dimensional feature generation module (310), and the multi-dimensional feature generation module (310) includes a time feature extraction submodule (311), a product feature extraction submodule (312), a user feature extraction submodule (313) and an environment feature extraction submodule (314), wherein: The time feature extraction submodule (311) is used to extract prediction features from time series data and extract periodic features through a time series period decomposition algorithm. , based on the quantification of the repeated fluctuation pattern of data within a fixed period; extracting trend features through sliding window linear fitting method , based on the change slope of the data in the window to reflect the long-term change direction; the mutation characteristics are extracted by the adjacent window difference threshold method , identify significant time nodes that deviate from the trend; The product feature extraction submodule (312) is used to extract prediction features from technical parameters and extract configuration difference features by calculating the Euclidean distance of parameter vectors. , quantify the differences in hardware parameters of different models; extract performance matching features through parameter combination synergy analysis ,The association relationship based on key parameters reflects the ,adaptability of the product to the usage scenario; The user feature extraction submodule (313) is used to extract prediction features from the operation trajectory and extract preference stability features by calculating the entropy value of the operation sequence. , quantify the consistency of users' long-term behavioral habits; extract behavioral conversion characteristics through key path conversion rate analysis , based on the user's operation chain from browsing to ordering, it reflects the stage characteristics of the decision-making process; The environmental feature extraction submodule (314) is used to extract prediction features from environmental data and generate regional correlation features by mining the co-occurrence patterns of meteorological, traffic data and historical sales. , quantify the impact of environmental factors on sales in different regions.
5. The automobile sales multi-dimensional prediction model construction system according to claim 4 is characterized in that: The dynamic feature engineering unit (300) further includes a dynamic screening mechanism module (320), wherein the dynamic screening mechanism module (320) includes a feature importance quantification submodule (321), a real-time contribution feedback submodule (322), a window iterative update submodule (323) and a feature set stability verification submodule (324), wherein: The feature importance quantification submodule (321) is used to calculate the influence of the prediction features extracted by the multi-dimensional feature generation module (310) on the prediction results. ,It is determined by comparing the difference in prediction error after including the prediction feature and randomly replacing the prediction feature value.,The influence value is positively correlated with the error change amplitude; The real-time contribution feedback submodule (322) is used to monitor the actual role of various prediction features in recent predictions. , based on the weighted cumulative value calculation of the feature importance within the sliding window, the weight coefficient of the recent data in the window is higher than that of the long-term data to ensure the timeliness of feedback; The window iteration update submodule (323) is used to dynamically adjust the feature set, when the influence of any prediction feature extracted by the multi-dimensional feature generation module (310) is Two consecutive windows are below the preset threshold, or their contribution When three consecutive windows show a monotonically decreasing trend, the prediction feature is automatically removed from the current feature set; at the same time, a new prediction feature is generated by combining cross-dimensional features. When the initial influence of the new prediction feature is When it exceeds the average level of the current feature set, it is included in the feature set to form an updated feature set; The feature set stability check submodule (324) is used to verify the validity of the updated feature set, and the overlap rate between the current feature set before the update and the feature set generated after the update is started by the calculation window iteration update submodule (323) and the forecast error fluctuation value To evaluate, when the overlap ratio Lower than the preset ratio or error fluctuation value When the threshold is exceeded, the multi-dimensional feature generation module (310) is triggered to re-extract the prediction features to optimize the feature set.
6. The automobile sales multi-dimensional prediction model construction system according to claim 5 is characterized in that: The multi-model collaborative prediction unit (400) includes a time series feature processing module (410) and a non-time series feature processing module (420), wherein: The time series feature processing module (410) is used to process time series features, and includes a recurrent neural network submodule (411) and an attention mechanism submodule (412), wherein: The recurrent neural network submodule (411) adopts a long short-term memory network structure and uses the periodic features generated by the time feature extraction submodule (311) , trend characteristics and mutation characteristics It takes the input as input, captures the dependencies of different time scales through the gating mechanism of input gate, forget gate and output gate, and outputs the hidden layer representation sequence of temporal features; The attention mechanism submodule (412) is used to enhance the influence weight of key temporal features and calculate the attention weight of each time step based on the hidden layer representation sequence. , for mutation features The time step and the recent time step in the prediction window are given higher weights, and the hidden layer representation sequence is converted into an aggregate vector of time series features through weighted summation as the intermediate prediction result of the time series channel; The non-temporal feature processing module (420) is used to process product attributes and user behavior features, and includes a tree-structured ensemble learning submodule (421) and a feature interaction submodule (422), wherein: The tree-like ensemble learning submodule (421) uses a gradient boosting tree algorithm to extract the configuration difference features generated by the product feature extraction submodule (312). , performance matching characteristics , and the preference stable features generated by the user feature extraction submodule (313) , behavioral conversion characteristics As input, a nonlinear prediction model is constructed through iterative training of multiple decision trees, and the basic prediction value of non-time series features is output; The feature interaction submodule (422) is used to mine cross-dimensional feature associations, using the regional association features generated by the environmental feature extraction submodule (314) As a link, calculate the interaction strength between product features and user features , generating high-order interaction features, and inputting the high-order interaction features into the tree-like ensemble learning submodule (421) for secondary training, and correcting the basic prediction value to form the final output of the non-temporal channel.
7. The automobile sales multi-dimensional prediction model construction system according to claim 6 is characterized in that: The multi-model collaborative prediction unit (400) further includes a model fusion module (430), wherein the model fusion module (430) includes an adaptive weight adjustment submodule (431), a prediction deviation monitoring submodule (432), and a fusion result verification submodule (433), wherein: The adaptive weight adjustment submodule (431) is used to dynamically fuse the output results of the time series feature processing module (410) and the non-time series feature processing module (420), and calculate the weight coefficient based on the real-time prediction deviation. : Time series model weight and non-series model weights The sum is 1, the weight value is generated by the inverse deviation function, and the weight update cycle is consistent with the prediction cycle; The prediction deviation monitoring submodule (432) is used to quantify the submodule output error and calculate the time series deviation of the intermediate prediction result of the time series feature processing module (410) based on the actual sales data. and the non-time-series deviation finally output by the non-time-series feature processing module (420) , the deviation sequence is stored through a sliding window with a length of 5 prediction periods to provide a historical basis for weight adjustment; The fusion result verification submodule (433) is used to verify the reliability of the fusion prediction by calculating the mean absolute error between the fusion result and the actual sales volume. And the improvement rate of the fusion result relative to the output of the time series feature processing module (410) and the non-time series feature processing module (420) Assessment: When Exceeds the preset threshold or When the adaptive weight adjustment submodule (431) is triggered to recalculate the weight, and the deviation sensitive features are fed back to the time series feature processing module (410) and the non-time series feature processing module (420), driving the secondary optimization of the feature weights.
8. The automobile sales multi-dimensional prediction model construction system according to claim 7 is characterized in that: The model training optimization unit (500) includes a data set time series partitioning module (510), a sub-model parameter iteration module (520), a training termination and adjustment module (530) and a training effect verification module (540), wherein: The data set time series partitioning module (510) is used to partition the training data according to the time series characteristics, and to partition the data set into a training set, a validation set, and a test set in chronological order based on the historical timestamps corresponding to the prediction features output by the dynamic feature engineering unit (300); when partitioning, the training set is ensured to cover at least 2 complete The corresponding historical data, validation set, and test set each cover at least 1 complete Corresponding historical data, while avoiding mutation characteristics Concentrated time intervals to avoid data distribution shifts that affect training stability; The sub-model parameter iteration module (520) is used to execute the sub-model core parameter update, and uses the gradient descent optimizer to optimize the two types of sub-models of the multi-model collaborative prediction unit (400) respectively: for the recurrent neural network with attention mechanism of the time series feature processing module (410), its network connection weight, attention weight and so on are optimized. Calculating coefficients and loop layer gate bias; optimizing the decision tree feature splitting threshold and node splitting gain coefficient of the tree-like ensemble learning model of the non-time-series feature processing module (420); dynamically setting the initial learning rate based on the complexity of the sub-model, decaying at a fixed ratio for each preset training round, and synchronously recording the error change of the validation set during the decay process; The training termination and adjustment module (530) is used to trigger the termination of training and handle abnormal training states: the preset termination condition is "the decrease in the error of the validation set for 3-5 consecutive rounds is lower than the preset error threshold", and the training is terminated when the condition is met; if the overfitting state of "the error of the training set continues to decrease while the error of the validation set increases" occurs, the time series sub-model is adjusted by increasing the dropout probability, and the non-time series sub-model is adjusted by increasing the tree pruning intensity; if the errors of the training set and the validation set do not improve for multiple consecutive rounds, less than or equal to 20% of recent data are selected from the test set to supplement the training set, and the parameter iteration process is restarted; The training effect verification module (540) is used to verify the effectiveness of the trained sub-model, using the test set that does not participate in parameter iteration as input to calculate the mean absolute percentage error between the predicted value and the actual sales volume. , goodness of fit ;when and When the preset model performance threshold is met, the training is determined to be effective and the final sub-model parameters are output; if the threshold requirement is not met, the module returns to the data set time series partitioning module (510) to readjust the partition ratio and execute the training process again.
9. The automobile sales multi-dimensional prediction model construction system according to claim 8, characterized in that: The prediction result output unit (600) includes a prediction result generation module (610), a visualization presentation module (620), a structured export module (630) and a historical backtracking comparison module (640), wherein: The prediction result generation module (610) is used to generate sales forecast data, receive the trained model parameters output by the model training optimization unit (500), load the real-time features output by the dynamic feature engineering unit (300), calculate and output the sales forecast value of the corresponding period according to a preset prediction period, and calculate the confidence interval corresponding to the prediction value based on the model prediction error characteristics; The visualization presentation module (620) is used to display forecast-related information, and synchronously presents sales forecast values, corresponding confidence intervals, and the influence weights of core forecast features through an interface, wherein the forecast values and confidence intervals are presented in the form of trend graphs, and the influence weights of core features are quantitatively presented in concise charts; The structured export module (630) is used to output the forecast result file, and supports exporting the sales forecast value, confidence interval and core feature influence weight in a standardized structured format. During the export process, the integrity of key data fields is automatically verified to ensure that there are no missing data. The historical backtracking comparison module (640) is used to compare the forecast with the historical data, call the historical sales data stored in the data acquisition unit (100), align the forecast result with the historical actual sales data according to the time dimension corresponding to the forecast period, calculate the deviation value between the two, and present the backtracking result in the form of a table or a comparison chart.
10. A method for constructing a multi-dimensional prediction model for automobile sales, based on the system for constructing a multi-dimensional prediction model for automobile sales according to any one of claims 1 to 9, characterized in that: The steps include: S100, Multi-source vehicle raw data collection: Connecting to the vehicle production database, sales terminal recording system, user behavior sensing equipment and government data disclosure interface through interface protocol adaptation; Process the high-frequency raw data of user behavior sensor devices through edge data preprocessing; store data through classified storage partitions according to data characteristics; Dynamically configure collection strategies through collection scheduling; verify data transmission and storage integrity through integrity checks, with check failures triggering retransmissions. S200, Data Cleaning and Standardization: Missing items in the original data are processed through missing value filling. Continuous numerical data uses linear interpolation, and discrete categorical data uses the weighted majority method. Outliers are detected and corrected through outlier identification and correction. Normally distributed data is identified based on 3 times the standard deviation, and non-normally distributed data is identified using the interquartile range method. Outliers are corrected using local weighted regression. Data standardization and structural transformation are achieved through data conversion. Numerical data uses min-max standardization, categorical data uses one-hot encoding, and unstructured data uses the bag-of-words model. The validity of the processed data is verified through preprocessing verification. S300, dynamic feature set generation: extracting multi-dimensional prediction features through multi-dimensional feature generation; Through the dynamic screening mechanism, features are screened, the feature influence and real-time contribution are calculated, and the feature set is iteratively updated. The validity of the updated feature set is verified through the feature set stability check, and features are re-extracted if they do not meet the requirements. S400, Dual-channel Collaborative Prediction and Fusion: Performs temporal feature processing on time series features, employs recurrent neural networks to capture time-scale dependencies, and incorporates attention mechanisms to enhance the impact of key temporal features. It also performs non-temporal feature processing on product attributes and user behavior features, constructs a prediction model using tree-structured ensemble learning, and mines cross-dimensional feature associations through feature interaction. It also fuses the outputs of the two channels, calculates adaptive weights, monitors prediction deviations, verifies the fusion results, and optimizes them. S500, model training optimization: divide the training set, validation set, and test set into time series by data set time series, ensuring that the complete cycle data is covered and avoiding the intervals with concentrated mutation features; The sub-model parameters are updated through iteration using a gradient descent optimizer, with the initial learning rate dynamically set and decayed per round. Training termination is triggered through training termination and adjustment to handle overfitting and training stagnation. The effectiveness of the trained sub-model is verified through training effect verification. If the requirements are not met, the dataset division is readjusted and training is resumed. S600, forecast result output: Generate sales forecast value and corresponding confidence interval based on the forecast result; display forecast value, confidence interval and core feature influence weight through visualization; Export forecast results in a standardized format through structured export and verify data integrity; align forecast results with historical sales data through historical backtesting, calculate deviations and present comparison results.
Citation Information
Patent Citations
A Multi-Scale Information Big Data Prediction Method for Automobile Sales Based on Attention Mechanism
CN113962750B
New energy automobile sales volume prediction method and system based on fusion model, and medium
CN118096243A
Sales prediction method and system based on prophet model and big data
CN113962745A
Automobile sales auxiliary decision-making method based on large model
CN120181972A
Multi-dimensional data statistics intelligent analysis and prediction system and method
CN120218985A
Cited By
Meteorological-power collaborative prediction method and device
CN121031917A
Multi-target mathematical calculation model construction method and application system
CN121388373A
A multi-target mathematical calculation model construction method and application system
CN121388373B
Method and system for constructing multi-source time series data fusion prediction model in BI analysis
CN121480859A
Logistics demand prediction method and system based on artificial intelligence
CN121961135A