Association factor prediction model training method, data completion method and equipment

By constructing a correlation factor prediction model and combining it with economic data using a decision tree ensemble model, the problem of missing automobile sales data was solved, the accuracy of data completion was improved, and high-quality market analysis and trend judgment were supported.

CN121901583APending Publication Date: 2026-04-21CHONGQING SOKON IND GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING SOKON IND GRP CO LTD
Filing Date
2025-12-05
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, automobile sales data is often incomplete due to reasons such as omissions in data statistics, loss of data records, or restrictions on data source statistical permissions. Traditional data completion methods are not very accurate, which affects the accuracy of market analysis and decision-making.

Method used

By acquiring manufacturer sales data and terminal sales data, a correlation factor prediction model is constructed, trained, and then a decision tree ensemble model is used in conjunction with economic data to predict the missing terminal sales data.

Benefits of technology

It improves the accuracy of missing data, provides high-quality complete datasets, and supports market analysis and trend judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901583A_ABST
    Figure CN121901583A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of data processing, in particular to a correlation factor prediction model training method, a data completion method and equipment. The method mainly comprises the following steps: acquiring existing manufacturer sales volume data, terminal sales volume data and economic data; obtaining an association factor set according to the manufacturer sales volume data and the terminal sales volume data; and inputting the association factor set, the manufacturer sales volume data and the economic data as training samples into a preset initial model for training to obtain an association factor prediction model. Wherein the manufacturer sales volume data and the terminal sales volume data are automobile sales volume data counted based on different data sources, and the association factors in the association factor set are used for representing the proportional relation of the manufacturer sales volume data and the terminal sales volume data corresponding to a single statistical period. By adopting the method, the accuracy of missing data complementation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method for training a correlation factor prediction model, a data completion method, and an apparatus. Background Technology

[0002] As a pillar industry of the national economy, the automotive industry's sales data is not only a core basis for automakers to formulate production plans and adjust marketing strategies, but also a key reference for policymakers to control market conditions and judge industry development trends. Currently, there are multiple statistical dimensions for automobile sales, and different data sources correspond to different types of sales statistics. For example, there are wholesale sales statistics from the China Association of Automobile Manufacturers (CAAM) and the China Passenger Car Association (CPCA), insurance registration sales statistics from the China Automotive Technology and Research Center (CATARC), and vehicle registration sales statistics from the Traffic Management Bureau of the Ministry of Public Security. These data reflect the circulation status of the automobile market from different perspectives.

[0003] In practical data applications, various sales data often suffer from missing information due to statistical omissions, lost records, or limitations in data source access permissions. To address this, techniques such as mean imputation and median imputation are commonly used to complete the data and make it as close to the actual situation as possible. However, these methods still suffer from low accuracy in data completion, which in turn affects the accuracy of subsequent work such as market analysis and trend judgment based on multi-source sales data, creating difficulties for decision-making by stakeholders in the automotive industry. Summary of the Invention

[0004] Based on this, it is necessary to provide a correlation factor prediction model training, data completion method, apparatus, computer equipment, and storage medium that are conducive to improving the accuracy of missing data completion, addressing at least one of the above-mentioned technical problems.

[0005] In a first aspect, embodiments of this disclosure provide a method for training a correlation factor prediction model. The model training method may include the following steps: acquiring existing manufacturer sales data, terminal sales data, and economic data; obtaining a set of correlation factors based on the manufacturer sales data and terminal sales data; and inputting the set of correlation factors, manufacturer sales data, and economic data as training samples into a preset initial model for training to obtain a correlation factor prediction model.

[0006] Among them, manufacturer sales data and terminal sales data are automobile sales data based on statistics from different data sources. The correlation factors in the correlation factor set are used to represent the proportional relationship between manufacturer sales data and terminal sales data for a single statistical period.

[0007] The correlation factor prediction model is used to obtain the target correlation factor for solving the missing terminal sales data of the specified period based on the manufacturer's sales data and economic data corresponding to the specified period.

[0008] In some embodiments, manufacturer sales data includes total manufacturer sales data and manufacturer vehicle sales data, terminal sales data includes total terminal sales data and terminal vehicle sales data, and the set of correlation factors includes a set of total sales coefficients and a set of vehicle coefficients.

[0009] Based on manufacturer sales data and end-user sales data, a set of correlation factors can be obtained, which may include the following steps: The ratio of total terminal sales data to total manufacturer sales data is used as the total sales coefficient to obtain the set of total sales coefficients. The ratio of terminal vehicle sales data to manufacturer vehicle sales data is used as the first candidate vehicle coefficient, thus obtaining the first candidate coefficient set. The ratio of the proportion of terminal vehicle models to the proportion of manufacturer vehicle models is used as the coefficient of the second candidate vehicle model, thus obtaining the set of second candidate coefficients. Based on preset evaluation indicators, the vehicle model coefficient set is determined from the first candidate coefficient set and the second candidate coefficient set.

[0010] Among them, the terminal vehicle proportion data is the ratio of terminal vehicle sales data to total terminal sales data, and the manufacturer vehicle proportion data is the ratio of manufacturer vehicle sales data to total manufacturer sales data.

[0011] The preset evaluation index is used to evaluate the dispersion of the first candidate coefficient set and the second candidate coefficient set. In some embodiments, the preset evaluation metrics include variance, standard deviation, coefficient of variation, range, and / or interquartile range.

[0012] In some embodiments, the correlation factor set, manufacturer sales data, and economic data are used as training samples and input into a preset initial model for training to obtain a correlation factor prediction model, which may include the following steps: The total sales coefficient set, vehicle model coefficient set, total sales data of manufacturers, sales data of vehicle models of manufacturers, vehicle model proportion data of manufacturers, and economic data are used as training samples and input into the preset initial model for training to obtain the correlation factor prediction model.

[0013] In some embodiments, the initial model is a decision tree ensemble model.

[0014] Using the set of correlation factors, manufacturer sales data, and economic data as training samples, and inputting them into a pre-defined initial model for training, a correlation factor prediction model can be obtained, which may include the following steps: Manufacturer sales data and economic data are combined into multiple feature matrix data corresponding to a single data statistical period, and the correlation factors are used as label data to obtain a labeled training dataset. The training dataset is input into the decision tree ensemble model for training, resulting in an association factor prediction model.

[0015] In some embodiments, manufacturer sales data are sales data from the China Association of Automobile Manufacturers (CAAM), wholesale data from the China Passenger Car Association (CPCA), or vehicle shipment data from automakers, while terminal sales data are insurance registration data, vehicle registration data, or retail data from the CPCA.

[0016] In a second aspect, embodiments of this disclosure provide a method for completing automobile sales data. This method may include the following steps: inputting the obtained manufacturer sales data and economic data for a specified period into a preset correlation factor prediction model to obtain a target correlation factor; and obtaining the missing terminal sales data for the specified period based on the target correlation factor and the manufacturer sales data.

[0017] Among them, the manufacturer sales data and the terminal sales data are automobile sales data statistically based on different data sources, and the correlation factor prediction model is obtained based on the model training method provided in any embodiment of the first aspect of this disclosure.

[0018] In some embodiments, manufacturer sales data includes total manufacturer sales data and manufacturer vehicle sales data, terminal sales data includes total terminal sales data and terminal vehicle sales data, and the target correlation factor includes total sales coefficient and vehicle coefficient.

[0019] Based on the target correlation factor and manufacturer sales data, the missing terminal sales data for a specified period can be obtained by the following steps: obtaining the missing total terminal sales data for a specified period based on the total sales coefficient and manufacturer total sales data; and obtaining the missing terminal model sales data for a specified period based on the model coefficient and manufacturer model sales data.

[0020] In some embodiments, manufacturer vehicle sales data includes manufacturer first data and manufacturer second data corresponding to different vehicle models.

[0021] Terminal vehicle sales data includes first-level terminal data and second-level terminal data corresponding to different vehicle models.

[0022] The vehicle model coefficient includes a first coefficient and a second coefficient corresponding to different vehicle models.

[0023] Based on vehicle model coefficients and manufacturer vehicle sales data, obtaining missing terminal vehicle sales data for a specified period can include the following steps: Based on the first coefficient and the manufacturer's first data, the first estimated data is obtained; Based on the second coefficient and the manufacturer's second data, the second estimated data is obtained; The second remaining data is determined based on the total terminal sales data and the first estimated data; When the second remaining data is greater than or equal to the second estimated data, the first estimated data is used as the first data of the terminal, and the second remaining data is used as the second data of the terminal. When the second remaining data is less than the second estimated data, the first estimated data is adjusted according to the difference between the second remaining data and the second estimated data to obtain the terminal first data, and the second estimated data is used as the terminal second data.

[0024] In a third aspect, embodiments of this disclosure provide a method for predicting terminal sales data, which may include the following steps: Terminal sales data, manufacturer sales data, and economic data corresponding to multiple historical statistical periods are used as training samples for the model. They are then input into a preset time series prediction model for training to obtain a terminal sales data prediction model. By inputting the manufacturer's sales data and economic data within the target statistical period corresponding to the period to be predicted into the terminal sales data prediction model, the terminal sales data corresponding to the period to be predicted can be obtained.

[0025] Among them, the period to be predicted is the future statistical period, the target statistical period corresponding to the period to be predicted is the historical statistical period, the number of periods of the target statistical period corresponds to the number of periods of the training samples in the training process of the terminal sales prediction model, and the terminal sales data corresponding to multiple historical statistical periods are the complete terminal sales data obtained after data completion using the data completion method provided in any embodiment of the second aspect.

[0026] Fourthly, embodiments of this disclosure provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the association factor prediction model training method provided in any embodiment of the first aspect of this disclosure.

[0027] In a fifth aspect, embodiments of the present disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the association factor prediction model training method provided in any embodiment of the first aspect of the present disclosure.

[0028] In a sixth aspect, embodiments of the present disclosure provide a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the data completion method provided in any embodiment of the second aspect of the present disclosure.

[0029] The aforementioned correlation factor prediction model training method, automobile sales data completion method, and computer equipment, through correlation factors representing the proportional relationship between manufacturer sales data and terminal sales data, are used to train a correlation factor prediction model in conjunction with economic data and manufacturer sales data. This correlation factor prediction model is used to obtain the target correlation factor for solving the missing terminal sales data for a specified period based on complete manufacturer sales data and economic data corresponding to that period. Compared with traditional data completion methods such as mean imputation and median imputation, this correlation factor prediction model fully combines the correlation patterns of automobile sales with the influence of the economic environment, improving the accuracy of missing data completion and providing a high-quality complete dataset for market analysis, trend research, and other work using the completed data. Attached Figure Description

[0030] Figure 1 This is a diagram illustrating the application environment of the correlation factor prediction model training method in some embodiments; Figure 2 This is a flowchart illustrating the training method for the correlation factor prediction model in some embodiments; Figure 3 This is a flowchart illustrating the steps involved in determining the set of associated factors in some embodiments; Figure 4 This is a flowchart illustrating the method for completing automobile sales data in some embodiments; Figure 5 This is a flowchart illustrating the steps involved in determining terminal sales data in some embodiments; Figure 6 This is a flowchart illustrating the steps involved in determining terminal vehicle sales data in some embodiments; Figure 7 This is a diagram showing the internal structure of a computer device in some embodiments. Detailed Implementation

[0031] To make the technical solutions and advantages of this disclosure clearer, the embodiments and related technical content of this disclosure will be further described in detail below with reference to the accompanying drawings and text description. It should be understood that the embodiments described below are only used to explain the technical solutions of the embodiments of this disclosure and are not intended to limit more possible implementations of this disclosure.

[0032] It should be noted that relational terms such as "first" and "second" appearing in this document are used only to distinguish things, states, or actions, and do not necessarily indicate or imply relative importance or order. The terms "including," "comprising," or any other variations thereof are used to indicate non-exclusive inclusion, and the included objects may not be limited to those listed in this document. The terms "multiple" or other variations are used to indicate that the number of objects is two or more.

[0033] In a first aspect, embodiments of this disclosure provide a method for training an association factor prediction model. This method can be applied to, for example... Figure 1 In the application environment shown, server 101 can communicate with terminal 102 via a network. Terminal 102 stores manufacturer sales data, terminal sales data, and economic data. Server 101 can be implemented using a standalone server or a server cluster consisting of multiple servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices.

[0034] Of course, the association factor prediction model training method provided in this disclosure can also be applied to more scenarios not shown in the figures.

[0035] Applying the correlation factor prediction model training method to Figure 1 Taking server 101 as an example, in some embodiments, such as Figure 2 As shown, the training method for the correlation factor prediction model includes steps S201 to S203 that can be executed by server 101. Each step is described in detail below.

[0036] Step S201: Obtain existing manufacturer sales data, terminal sales data, and economic data.

[0037] Among them, manufacturer sales data and terminal sales data are automobile sales data based on statistics from different data sources.

[0038] Manufacturer sales data can be based on sales figures from wholesale vehicle quantities collected by automakers. Retail sales data can be based on sales figures from vehicles ultimately delivered to consumers or entering actual use. Economic data can be macroeconomic data published on national information disclosure websites, such as GDP (Gross Domestic Product) data, disposable income data, and automotive industry policy index data. Specifically, GDP data can include total GDP, GDP growth data, and GDP per capita.

[0039] In some specific examples, manufacturer sales data can be wholesale automobile sales data compiled by the China Association of Automobile Manufacturers (CAAM) or the China Passenger Car Association (CPCA) (referred to as CAAM sales data or CPCA sales data), while terminal sales data can be automobile insurance data or vehicle registration data.

[0040] Manufacturer sales data, terminal sales data, and economic data can be in Excel (a proprietary software name defined by Microsoft) format or CSV (a plain text format with different data fields separated by commas) format.

[0041] Existing manufacturer sales data, retail sales data, and economic data can be based on an annual, semi-annual, quarterly, or monthly statistical period. They can be manufacturer sales data, retail sales data, and economic data corresponding to continuous statistical periods, or they can be manufacturer sales data, retail sales data, and economic data corresponding to multiple fragmented statistical periods.

[0042] Specifically, existing manufacturer sales data, terminal sales data, and economic data can be obtained from a pre-built storage device on terminal 102, or from a publicly available data platform on server 101. For example, if manufacturer sales data is from the China Association of Automobile Manufacturers (CAAM) and terminal sales data is from insurance registration data, manufacturer sales data can be obtained from CAAM's official website, while terminal sales data can be obtained from the national vehicle insurance registration platform or other official channels.

[0043] Step S202: Obtain the set of correlation factors based on manufacturer sales data and terminal sales data.

[0044] The correlation factors in the correlation factor set are used to represent the proportional relationship between manufacturer sales data and terminal sales data for a single statistical period. That is, there are corresponding correlation factors for manufacturer sales data and terminal sales data for each statistical period.

[0045] In some specific examples, the correlation factor can be the ratio of terminal sales data to manufacturer sales data.

[0046] Specifically, based on the manufacturer sales data and terminal sales data corresponding to each statistical period, the correlation factors corresponding to each statistical period are obtained, resulting in a set of correlation factors.

[0047] Step S203: Use the set of correlation factors, manufacturer sales data, and economic data as training samples, input them into the preset initial model for training, and obtain the correlation factor prediction model.

[0048] Among them, the correlation factor prediction model is used to obtain the target correlation factor for solving the missing terminal sales data of the specified period based on the manufacturer's sales data and economic data corresponding to the specified period.

[0049] The specified period refers to the statistical period corresponding to the missing terminal sales data. Based on this correlation factor prediction model, when there is a missing period in the terminal sales data, the manufacturer's sales data and economic data corresponding to the missing period (i.e., the specified period) are input into the correlation factor prediction model to obtain the target correlation factor corresponding to the specified period, and thus obtain the missing terminal sales data for the specified period. It is easy to understand that the manufacturer's sales data and economic data corresponding to the specified period here are both valid and complete data.

[0050] Step S203 can involve using manufacturer sales data and economic data as input features, and correlation factors as target labels, to train an initial model and iteratively optimize it, resulting in a correlation factor prediction model capable of capturing the intrinsic correlation between manufacturer sales data, economic data, and correlation factors. Specifically, the training samples can include multiple sets of sample data, each set corresponding to a statistical period. That is, the manufacturer sales data and economic data corresponding to a statistical period, along with the correlation factors corresponding to that period, constitute a set of sample data. For example, manufacturer sales data for the first quarter of 2023, economic data for the first quarter of 2023, and correlation factors for the first quarter of 2023 form a correspondence between features and labels, serving as a set of sample data.

[0051] Taking the manufacturer's sales data as the China Association of Automobile Manufacturers' sales data and the terminal sales data as the insurance registration data as an example, in some specific scenarios, the insurance registration data can only be obtained from 2015 to 2024, and the data from 2005 to 2014 is missing. Correspondingly, the specified period is: 10 statistical periods from 2005 to 2014.

[0052] Step S201 may include: obtaining manufacturer sales data, terminal sales data, and economic data for 10 statistical periods between 2015 and 2024.

[0053] Step S202 may include: obtaining 10 correlation factors corresponding to 10 statistical periods from 2015 to 2024 based on manufacturer sales data and terminal sales data from 2015 to 2024, and obtaining a set of correlation factors.

[0054] The correlation factor model obtained in step S203 can be used to predict 10 target correlation factors corresponding to a specified period based on manufacturer sales data and economic data from 2005 to 2014. Then, based on the target correlation factors and manufacturer sales data from 2005 to 2014, the missing terminal sales data for the specified period can be obtained.

[0055] In the aforementioned correlation factor prediction model training method, a correlation factor prediction model is trained by combining economic data and manufacturer sales data, using correlation factors representing the proportional relationship between manufacturer sales data and terminal sales data. This model is used to obtain target correlation factors for solving missing terminal sales data for a specified period, based on complete manufacturer sales data and economic data. Compared to traditional data completion methods such as mean imputation and median imputation, this correlation factor prediction model fully integrates the correlation patterns of automobile sales with the influence of the economic environment, improving the accuracy of missing data completion. It can provide a high-quality, complete dataset for market analysis, trend research, and other work using completed data.

[0056] In some embodiments, manufacturer sales data includes total manufacturer sales data and manufacturer vehicle sales data. Terminal sales data includes total terminal sales data and terminal vehicle sales data. The set of correlation factors includes a set of coefficients related to total sales and a set of coefficients related to vehicle models.

[0057] In other words, the statistical dimensions of automobile sales data include total sales and sales of different models categorized by vehicle type.

[0058] like Figure 3 As shown, step S202 may include steps S301 to S304.

[0059] Step S301: Use the ratio of total terminal sales data to total manufacturer sales data as the total sales coefficient to obtain a set of total sales coefficients.

[0060] Specifically, for each statistical period, the total sales coefficient is obtained based on the ratio of total terminal sales data to total manufacturer sales data, and then multiple total sales coefficients corresponding to multiple statistical periods are obtained, i.e., the set of total sales coefficients.

[0061] Step S302: Use the ratio of terminal vehicle sales data to manufacturer vehicle sales data as the first candidate vehicle coefficient to obtain the first candidate coefficient set.

[0062] Step S303: Use the ratio of the terminal vehicle proportion data to the manufacturer vehicle proportion data as the second candidate vehicle coefficient to obtain the second candidate coefficient set.

[0063] Among them, the terminal vehicle proportion data is the ratio of terminal vehicle sales data to total terminal sales data, and the manufacturer vehicle proportion data is the ratio of manufacturer vehicle sales data to total manufacturer sales data.

[0064] Step S304: Based on preset evaluation indicators, determine the vehicle model coefficient set from the first candidate coefficient set and the second candidate coefficient set.

[0065] The preset evaluation index is used to evaluate the dispersion of the first candidate coefficient set and the second candidate coefficient set.

[0066] Specifically, based on preset evaluation indicators, the set of candidate coefficients with smaller dispersion between the first and second candidate coefficient sets is determined as the vehicle model coefficient set.

[0067] It is easy to understand that the set of vehicle coefficients corresponding to different vehicle models can all be obtained based on the first candidate vehicle coefficient or the second candidate vehicle coefficient, or the set of vehicle coefficients for some vehicle models can be obtained based on the first candidate vehicle coefficient, and the set of vehicle coefficients for some vehicle models can be obtained based on the second candidate vehicle coefficient.

[0068] When the vehicle model coefficient is the first candidate vehicle model coefficient, the corresponding vehicle's terminal sales data can be determined based on the product of the vehicle model coefficient and the manufacturer's vehicle sales data. When the vehicle model coefficient is the second candidate vehicle model coefficient, the corresponding vehicle's terminal sales data can be determined based on the product of the total terminal sales data, the vehicle model coefficient, and the manufacturer's vehicle proportion data, where the manufacturer's vehicle proportion data is the ratio of the manufacturer's vehicle sales data to the manufacturer's total sales data.

[0069] In some specific scenarios, for example, both manufacturer vehicle sales data and terminal vehicle sales data include sales data corresponding to three different vehicle types: sedans, SUVs (Sport Utility Vehicles), and MPVs (Multi-Purpose Vehicles / Minivans), with the data statistics period being 10 statistical periods from 2015 to 2024.

[0070] Correspondingly, the vehicle model coefficient set can include the sedan vehicle model coefficient set, the SUV vehicle model coefficient set, and the MPV vehicle model coefficient set.

[0071] Step S302 may include the following steps: calculating the ratio of terminal vehicle sales data to manufacturer vehicle sales data for each different vehicle type in each statistical period (i.e., each year), to obtain a first candidate coefficient set for different vehicle types. That is, the first candidate coefficient set includes a first candidate coefficient set for sedans, a first candidate coefficient set for SUVs, and a first candidate coefficient set for MPVs.

[0072] Step S303 may include the following steps: calculating the terminal vehicle proportion data for each different car model in each year; calculating the manufacturer vehicle proportion data for each different car model in each year; and obtaining a second candidate coefficient set based on the ratio of the terminal vehicle proportion data to the manufacturer vehicle proportion data. The second candidate coefficient set includes a second candidate coefficient set for sedans, a second candidate coefficient set for SUVs, and a second candidate coefficient set for MPVs.

[0073] Step S304 may include the following steps: determining a set of car model coefficients from a first set of candidate coefficients for cars and a second set of candidate coefficients for cars based on preset evaluation indicators; determining a set of SUV model coefficients from a first set of candidate coefficients for SUVs and a second set of candidate coefficients for SUVs based on preset evaluation indicators; and determining a set of MPV model coefficients from a first set of candidate coefficients for MPVs and a second set of candidate coefficients for MPVs based on preset evaluation indicators.

[0074] By using preset evaluation indicators to select a set of candidate coefficients with lower dispersion as vehicle model coefficients, the system exhibits stronger time stability, solving the problems of large fluctuations and data distortion caused by traditional ratio indicators. This allows the correlation factor prediction model to more accurately capture the relationship between correlation factors and sales data, improving the accuracy of correlation factor prediction and further enhancing the accuracy of missing data completion.

[0075] In some embodiments, the preset evaluation metrics may include variance, standard deviation, coefficient of variation, range, and / or interquartile range.

[0076] Step S304 may include: calculating the variance values ​​of the first candidate coefficient set and the second candidate coefficient set respectively, and selecting the candidate coefficient set with the smaller variance value as the vehicle model coefficient set.

[0077] For example, if the variance of the first candidate coefficient set is calculated to be 0.0023 and the variance of the second candidate coefficient set is 0.0006, and the variance of the second candidate coefficient set is significantly smaller than that of the first candidate coefficient set, it indicates that the dispersion of the first candidate coefficient set is higher than that of the second candidate coefficient set. In other words, the stability of the second candidate vehicle coefficient is higher. Therefore, the second candidate vehicle coefficient is selected as the target vehicle coefficient, and the corresponding second candidate coefficient set is used as the vehicle coefficient set.

[0078] In some specific examples, evaluation metrics may include a combination of one or more of variance, standard deviation, coefficient of variation, range, and interquartile range.

[0079] Variance is a statistic that measures the degree of deviation of each data point in a dataset from the mean of the dataset. Variance is essentially the squared mean of the deviations. The larger the value, the more dispersed the data points are around the mean, i.e., the greater the dispersion. The smaller the value, the more concentrated the data points are, i.e., the smaller the dispersion.

[0080] Standard deviation is the square root of variance and is a core statistic for measuring the absolute fluctuation of data. The larger the variance, the larger the standard deviation, and the more volatile the data; conversely, the smaller the variance, the more moderate the fluctuation.

[0081] The coefficient of variation (COP) is the ratio of the standard deviation to the mean of a dataset. It is a statistic that measures the relative volatility of data and allows for direct comparison of the dispersion of different dimensions and means. A smaller COP indicates less relative volatility and better stability in the dataset, while a larger COP indicates greater volatility.

[0082] The range is the difference between the maximum and minimum values ​​in a dataset. It is a statistical measure of the extreme fluctuation range of data. The smaller the range, the better the stability of the dataset.

[0083] The interquartile range (ICM) is the difference between the upper and lower quartiles of a dataset, measuring the fluctuation range of the middle 50% of the core data. A smaller ICM indicates less fluctuation and greater stability in the core data, while a larger ICM indicates greater fluctuation.

[0084] Specifically, the evaluation indicators, which include a combination of multiple dimensions, may include the following first evaluation indicator, second evaluation indicator, and third evaluation indicator.

[0085] The primary evaluation indicators can be the standard deviation and the interquartile range. The standard deviation is used to evaluate the fluctuation of the coefficient, and the interquartile range is used to evaluate the stability of the coefficient.

[0086] The second evaluation index can be the coefficient of variation and the interquartile range. The coefficient of variation is used to compare the relative fluctuations between coefficients, and the interquartile range is used to evaluate the stability of the coefficients.

[0087] The third evaluation metric can be the range, standard deviation, and coefficient of variation. The range can be used to quickly define the fluctuation range, the standard deviation can show the absolute fluctuation, and the coefficient of variation can supplement the relative fluctuation, ensuring that the selected set of vehicle coefficients is a set of coefficients with better stability.

[0088] The selection of evaluation indicators can be set by those skilled in the art according to the actual situation, and no special restrictions are imposed here.

[0089] In some embodiments, the correlation factor prediction model is obtained by using the set of correlation factors, manufacturer sales data, and economic data as training samples and inputting them into a preset initial model. This may include the following steps: using the set of total sales coefficients, the set of vehicle model coefficients, the manufacturer's total sales data, the manufacturer's vehicle model sales data, the manufacturer's vehicle model proportion data, and economic data as training samples and inputting them into a preset initial model for training.

[0090] In some specific scenarios, when both manufacturer vehicle sales data and terminal vehicle sales data include sales data for three different vehicle types: sedans, SUVs, and MPVs, the training dataset for the correlation factor prediction model includes a set of total sales coefficients for multiple statistical periods, a set of coefficients for sedans, SUVs, and MPVs, manufacturer total sales data, manufacturer vehicle sales data (sales data for the three different vehicle types respectively), and economic data.

[0091] In some embodiments, the initial model is a decision tree ensemble model, specifically, it may be XGBoost (ExtremeGradient Boosting).

[0092] Step S203 may include the following steps: combining manufacturer sales data and economic data into multiple sets of feature matrix data corresponding to a single data statistical period, and using correlation factors as label data to obtain a labeled training dataset; inputting the training dataset into a decision tree ensemble model for training to obtain a correlation factor prediction model.

[0093] In some specific examples, the training parameters for a decision tree ensemble model may include the learning rate, maximum depth of the decision tree, number of weak learners, training sample sampling ratio, and feature sampling ratio. By optimizing the initial values ​​of the training parameters within a pre-defined optimization search range based on the training dataset, the optimized target training parameters are determined, and then the association factor prediction model is trained.

[0094] By adopting the XGBoost model to support grid search optimization of hyperparameters (such as the training parameters mentioned above), and by flexibly adjusting key parameters, the computational efficiency is improved while ensuring accuracy. Moreover, there is no need to query non-public data sources. Modeling can be completed solely by relying on publicly available data such as the China Association of Automobile Manufacturers' official website and the National Auto Insurance Registration Platform, which reduces the cost of data acquisition and the application threshold.

[0095] In some embodiments, manufacturer sales data may be sales data from the China Association of Automobile Manufacturers (CAAM), wholesale data from the China Passenger Car Association (CPCA), or vehicle shipment data from automakers; terminal sales data may be insurance registration data, vehicle registration data, or retail data from the CPCA.

[0096] Among them, the sales data from the China Association of Automobile Manufacturers (CAAM) is based on the sales data reported by automakers to their dealers. This data belongs to the statistical data of the production and distribution process.

[0097] The wholesale data from the China Passenger Car Association (CPCA) is compiled by the CPCA based on the sales data reported by automakers to their dealers.

[0098] Automaker shipment data refers to the number of vehicles that automakers transport from the factory to dealerships in various regions based on orders placed with dealers.

[0099] The insurance registration data is compiled by the China Automotive Technology and Research Center (CATARC) in conjunction with insurance departments or the national vehicle insurance registration platform, based on the compulsory traffic accident liability insurance information paid by consumers after purchasing a vehicle. It can accurately reflect the actual situation of vehicles being delivered to end consumers.

[0100] The vehicle registration data is compiled by the Traffic Management Bureau of the Ministry of Public Security, based on records of vehicle license plates issued by local vehicle management offices. This data is one of the core indicators reflecting the actual sales volume at the terminal level.

[0101] The retail data from the China Passenger Car Association (CPCA) is compiled by the CPCA based on actual vehicle sales data provided by vehicle dealers.

[0102] In some embodiments, the correlation factor prediction model training method may further include the following steps: after acquiring existing manufacturer sales data, terminal sales data and economic data, perform data preprocessing on the manufacturer sales data and terminal sales data.

[0103] Data preprocessing may include steps such as missing value checking, outlier handling, format conversion, and maximum / minimum value normalization.

[0104] Specifically, in some examples, outlier handling can employ the interquartile range (ICM) method to identify outliers. Manufacturer sales data with ICM values ​​within a preset outlier range are replaced using the manufacturer's moving average, and terminal sales data with ICM values ​​within the preset outlier range are replaced using the terminal's moving average. For example, if there are outliers in the 2020 manufacturer sales data, the average of the 2019 and 2021 manufacturer sales data is used for replacement.

[0105] By using the interquartile range method to identify outliers and replacing them with annual moving averages, combined with maximum and minimum value normalization, data quality is ensured while the preprocessing process is simplified.

[0106] In a second aspect, embodiments of this disclosure provide a method for completing automobile sales data, such as... Figure 4 As shown, the method for completing automobile sales data may include steps S401 and S402.

[0107] Step S401: Input the obtained manufacturer sales data and economic data corresponding to the specified period into the preset correlation factor prediction model to obtain the target correlation factor.

[0108] Step S402: Based on the target correlation factor and manufacturer sales data, obtain the terminal sales data missing for the specified period.

[0109] Among them, manufacturer sales data and terminal sales data are automobile sales data statistically derived from different data sources, and the correlation factor prediction model is obtained based on the correlation factor prediction model training method provided in any embodiment of the first aspect of this disclosure. The target correlation factor corresponds to a specified period. When the specified period includes multiple statistical periods, the target correlation factor also includes multiple factors, and the number is consistent with the number of statistical periods.

[0110] In some embodiments, manufacturer sales data includes total manufacturer sales data and manufacturer vehicle sales data. Terminal sales data includes total terminal sales data and terminal vehicle sales data. Target correlation factors include total sales coefficient and vehicle coefficient.

[0111] Correspondingly, such as Figure 5As shown, step S402 may include steps S501 and S502.

[0112] Step S501: Based on the total sales coefficient and the manufacturer's total sales data, obtain the terminal total sales data missing for the specified period.

[0113] Step S502: Based on the vehicle model coefficient and manufacturer vehicle sales data, obtain the terminal vehicle sales data missing for the specified period.

[0114] In some embodiments, manufacturer vehicle sales data includes manufacturer first data and manufacturer second data corresponding to different vehicle models, and terminal vehicle sales data includes terminal first data and terminal second data corresponding to different vehicle models. Correspondingly, vehicle model coefficients include first coefficients and second coefficients corresponding to different vehicle models.

[0115] In some specific examples, different vehicle types can include sedans and SUVs; in other examples, different vehicle types can include sedans, SUVs, and MPVs; and of course, different vehicle types can also include many other different vehicle types.

[0116] like Figure 6 As shown, step S502 may include steps S601 to S605.

[0117] Step S601: Obtain the first estimated data based on the first coefficient and the manufacturer's first data.

[0118] Step S602: Obtain the second estimated data based on the second coefficient and the manufacturer's second data.

[0119] Step S603: Determine the second remaining data based on the total terminal sales data and the first estimated data.

[0120] Step S604: When the second remaining data is greater than or equal to the second estimated data, the first estimated data is used as the terminal first data, and the second remaining data is used as the terminal second data.

[0121] Step S605: When the second remaining data is less than the second estimated data, adjust the first estimated data according to the difference between the second remaining data and the second estimated data to obtain the terminal first data, and use the second estimated data as the terminal second data.

[0122] The process of adjusting the first estimated data based on the difference to obtain the first terminal data can be achieved by subtracting the difference from the first estimated data.

[0123] Taking sedans and SUVs as examples, steps S601 to S605 can be described as follows: Based on the sedan coefficient and manufacturer's sedan data, obtain the estimated sedan data; based on the SUV coefficient and manufacturer's SUV data, obtain the estimated SUV data; based on the total terminal sales data and sedan estimated data, determine the remaining SUV data; when the remaining SUV data is greater than or equal to the estimated SUV data, use the sedan estimated data as the terminal sedan data and the remaining SUV data as the terminal SUV data; when the remaining SUV data is less than the estimated SUV data, adjust the sedan estimated data based on the difference between the remaining SUV data and the estimated SUV data to obtain the terminal sedan data, and use the estimated SUV data as the terminal SUV data.

[0124] Specifically, the coefficients for sedans and SUVs, both determined based on the ratio of end-user vehicle proportion data to manufacturer vehicle proportion data, can be expressed by the following formula: ; ; ; like ≥ ,but As terminal car data, Data for end-user SUVs; like < According to and Adjust the difference between them Obtain the terminal car data, and Data from end-user SUVs.

[0125] in, Forecast data for cars, Car coefficient, For manufacturer's car data, Manufacturer's total sales data Total terminal sales data Forecast data for SUVs, For SUV coefficients, For SUV data from manufacturers, Remaining data for SUVs.

[0126] In some specific examples, manufacturer vehicle sales data includes manufacturer first data, manufacturer second data, and manufacturer third data corresponding to different vehicle models; terminal vehicle sales data includes terminal first data, terminal second data, and terminal third data corresponding to different vehicle models; and correspondingly, vehicle model coefficients include first coefficients, second coefficients, and third coefficients corresponding to different vehicle models.

[0127] Correspondingly, referring to steps S601 to S603 above, first estimated data, second estimated data, third estimated data, and third residual data are obtained. When the third residual data is greater than or equal to the third estimated data, the first estimated data is used as the terminal's first data, the second estimated data is used as the terminal's second data, and the third residual data is used as the terminal's third data. When the third residual data is less than the third estimated data, the first and second estimated data are adjusted proportionally according to the difference between the third residual data and the third estimated data to obtain the terminal's first data and terminal's second data, and the third estimated data is used as the terminal's second data.

[0128] Taking different vehicle types, including sedans, SUVs, and MPVs, as examples, and when the sedan coefficient and SUV coefficient are both second candidate vehicle coefficients determined based on the ratio of terminal vehicle proportion data to manufacturer vehicle proportion data, and the MPV coefficient is the first candidate vehicle coefficient determined based on the ratio of terminal vehicle sales data to manufacturer vehicle sales data, step S502 can be expressed by the following formula: ; ; ; ; like ≥ ,but As terminal car data As data for end-user SUVs, As terminal MPV data; like < According to and The difference between them is adjusted proportionally. and Obtain terminal sedan data and terminal SUV data, and As terminal MPV data.

[0129] in, Forecast data for cars, Car coefficient, For manufacturer's car data, Manufacturer's total sales data Total terminal sales data Forecast data for SUVs, For SUV coefficients, For SUV data from manufacturers, For the remaining data of SUVs, For MPV forecast data, MPV coefficient This is the remaining data for MPV.

[0130] By comparing the remaining data of some models with the estimated data, the final sales data of the models can be determined based on different comparison results. When the remaining data is less than the estimated data, the estimated data of other models can be adjusted proportionally according to the difference between the remaining data and the estimated data to obtain the final sales data of other models. This can ensure the rationality and completeness of the segmented model data and avoid the dimensional data deviation caused by a single completion logic.

[0131] Taking terminal sales data as insurance registration data as an example, by using the above-mentioned data completion method to complete the data, we can obtain insurance registration data for the complete data statistical period. This data can then be directly applied to core scenarios such as auto insurance risk factor analysis, differentiated rate setting, and claims risk screening. This provides reliable data references for insurance companies to accurately identify high-risk customer groups and optimize claims processes, and for automakers to carry out capacity planning and market strategy formulation. It solves the long-standing pain point of incomplete insurance registration data in the industry.

[0132] For more specific limitations on the method of completing automobile sales data, please refer to the limitations on the training method of the correlation factor prediction model mentioned above, which will not be repeated here.

[0133] It should be understood that, although Figures 2-6 The steps in the flowchart are shown sequentially according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Figures 2-6 Unless otherwise expressly stated herein, the steps illustrated and other steps involved in the embodiments are not subject to strict order restrictions and may be performed in other orders. Furthermore, at least some steps in the foregoing embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0134] In a third aspect, embodiments of this disclosure provide a method for predicting terminal sales data, which may include the following steps: Terminal sales data, manufacturer sales data, and economic data corresponding to multiple historical statistical periods are used as training samples for the model. They are then input into a preset time series prediction model for training to obtain a terminal sales data prediction model. By inputting the manufacturer's sales data and economic data within the target statistical period corresponding to the period to be predicted into the terminal sales data prediction model, the terminal sales data corresponding to the period to be predicted can be obtained.

[0135] Among them, the period to be predicted is the future statistical period, the target statistical period corresponding to the period to be predicted is the historical statistical period, the number of periods of the target statistical period corresponds to the number of periods of the training samples in the training process of the terminal sales prediction model, and the terminal sales data corresponding to multiple historical statistical periods are the complete terminal sales data obtained after data completion using the data completion method provided in any embodiment of the second aspect.

[0136] In some specific examples, after the terminal sales data is supplemented, complete terminal sales data for the statistical period is obtained. Based on the complete data, a terminal sales data prediction model can be constructed using time series modeling.

[0137] In some specific scenarios, the time series prediction model used can be a linear regression model, a random forest regression prediction model, an XGBoost model, an LSTM (Long Short-Term Memory) model, or other similar models.

[0138] For Random Forest and XGBoost, time-series features need to be pre-constructed, such as the mean and difference. For each sample, lagged features for each year are added. For example, when building a model to predict data for the next 3 years (the period to be predicted) using 8 years (the target statistical period), to predict terminal sales data from 2025 to 2027, complete terminal sales data, manufacturer sales data, and economic data from 2017 to 2024 need to be added as input to each row of samples.

[0139] The technical effects of the data completion method involved in the embodiments of this disclosure are demonstrated by comparing experimental data below.

[0140] A prediction model is trained using terminal sales data to be completed, terminal sales data completed by linear interpolation, terminal sales data completed by mean filling, and terminal sales data completed by the data completion method involved in the embodiments of this disclosure. Terminal sales data, manufacturer sales data, and economic data from 2014 to 2021 are used as input samples to predict terminal sales data from 2022 to 2024. Then, based on the terminal sales prediction data and the actual terminal sales data, the error rate data in the table below is determined.

[0141] The terminal sales data used in the comparative experiment were insurance registration data, and the manufacturer sales data were from the China Association of Automobile Manufacturers (CAAM). Official insurance registration data available only covers the period from 2015 to 2024 with complete records; the period from 2005 to 2014 has no data records and is missing. The insurance registration data from 2005 to 2014 was supplemented using linear interpolation, mean imputation, and the data completion method provided in any embodiment of the second aspect of this disclosure, and then a prediction model was constructed.

[0142]

[0143] The error rate is calculated as (predicted value - actual value) / actual value.

[0144] It's easy to understand that without data completion, the existing insurance registration data only includes 10 years of data. If we use 8 years of data to predict 3 years of data, there will be no complete training set. If we reduce the training set period, for example, using 5 years of data to predict 3 years of data, even with a rolling sample construction method, there will only be two samples in total: one for training and one for testing, leading to severe overfitting of the model. Using 3 years to predict 3 years will fail to capture long-term historical trends.

[0145] Based on the above experimental data analysis, it can be seen that the data completion method disclosed herein has high data completion accuracy. The prediction model trained on the terminal sales data after data completion by using the target correlation factor obtained by the correlation factor prediction model has better prediction accuracy and better prediction stability compared with the prediction model trained on the terminal sales data completed by other data completion methods, providing more reliable data support for scenarios such as car insurance rate determination and automobile market supply and demand prediction.

[0146] In a fourth aspect, embodiments of this disclosure provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the association factor prediction model training method provided in any embodiment of the first aspect of this disclosure.

[0147] In some embodiments, the computer device may be a server, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores manufacturer sales data, terminal sales data, and economic data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements the correlation factor prediction model training method in any embodiment of this document.

[0148] Those skilled in the art will understand that Figure 7 The structures shown are merely block diagrams of some structures related to the embodiments of this disclosure and do not constitute a limitation on the computer devices to which the embodiments of this disclosure are applied. Specific computer devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.

[0149] In a fifth aspect, embodiments of the present disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the association factor prediction model training method provided in any embodiment of the first aspect of the present disclosure.

[0150] The computer-readable storage medium may be Figure 7 The computer-readable storage medium in the computer device shown.

[0151] In a sixth aspect, embodiments of the present disclosure provide a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the data completion method provided in any embodiment of the second aspect of the present disclosure.

[0152] In some embodiments, the terminal device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the vehicle sales data completion method in any embodiment of this document. The display screen of the computer device can be a liquid crystal display screen or an e-ink display screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad provided on the casing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0153] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The aforementioned computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments of this disclosure can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0154] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this disclosure.

[0155] The above embodiments merely illustrate several implementation methods of this disclosure, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of protection of this disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the appended claims.

Claims

1. A method for training a correlation factor prediction model, characterized in that, The method includes: Acquire existing manufacturer sales data, terminal sales data, and economic data, wherein the manufacturer sales data and the terminal sales data are automobile sales data statistically analyzed based on different data sources; Based on the manufacturer sales data and the terminal sales data, a set of correlation factors is obtained. The correlation factors in the set of correlation factors are used to represent the proportional relationship between manufacturer sales data and terminal sales data for a single statistical period. The correlation factor set, the manufacturer sales data, and the economic data are used as training samples and input into a preset initial model for training to obtain the correlation factor prediction model. The correlation factor prediction model is used to obtain the target correlation factor for solving the missing terminal sales data of the specified period based on the manufacturer sales data and economic data corresponding to the specified period.

2. The method according to claim 1, characterized in that, The manufacturer sales data includes total manufacturer sales data and manufacturer model sales data; the terminal sales data includes total terminal sales data and terminal model sales data; the set of correlation factors includes a set of total sales coefficients and a set of model coefficients. The set of correlation factors obtained based on the manufacturer's sales data and the terminal sales data includes: The ratio of the total sales data of the terminals to the total sales data of the manufacturers is used as the total sales coefficient to obtain the set of total sales coefficients; The ratio of the terminal vehicle sales data to the manufacturer vehicle sales data is used as the first candidate vehicle coefficient, thus obtaining the first candidate coefficient set. The ratio of the proportion of terminal vehicle models to the proportion of manufacturer vehicle models is used as the second candidate vehicle model coefficient to obtain the second candidate coefficient set; wherein, the proportion of terminal vehicle models is the ratio of the sales volume of terminal vehicle models to the total sales volume of terminal vehicles, and the proportion of manufacturer vehicle models is the ratio of the sales volume of manufacturer vehicle models to the total sales volume of manufacturers. Based on preset evaluation indicators, the vehicle model coefficient set is determined from the first candidate coefficient set and the second candidate coefficient set. The preset evaluation indicators are used to evaluate the dispersion of the first candidate coefficient set and the second candidate coefficient set.

3. The method according to claim 2, characterized in that, The preset evaluation indicators include variance, standard deviation, coefficient of variation, range, and / or interquartile range.

4. The method according to claim 2, characterized in that, The step of using the set of correlation factors, the manufacturer sales data, and the economic data as training samples and inputting them into a preset initial model for training to obtain the correlation factor prediction model includes: The total sales coefficient set, the vehicle model coefficient set, the manufacturer's total sales data, the manufacturer's vehicle model sales data, the manufacturer's vehicle model proportion data, and the economic data are used as training samples and input into a preset initial model for training to obtain the correlation factor prediction model.

5. The method according to claim 1, characterized in that, The initial model is a decision tree ensemble model; The step of using the set of correlation factors, the manufacturer sales data, and the economic data as training samples and inputting them into a preset initial model for training to obtain the correlation factor prediction model includes: The manufacturer sales data and the economic data are combined into multiple feature matrix data corresponding to a single data statistical period, and the correlation factors are used as label data to obtain a labeled training dataset. The training dataset is input into the decision tree ensemble model for training to obtain the association factor prediction model.

6. The method according to claim 1, characterized in that, The manufacturer sales data refers to sales data from the China Association of Automobile Manufacturers (CAAM), wholesale data from the China Passenger Car Association (CPCA), or vehicle shipment data from automakers. The terminal sales data refers to insurance registration data, vehicle registration data, or retail data from the CPCA.

7. A method for completing automobile sales data, characterized in that, The method includes: The obtained manufacturer sales data and economic data for the specified period are input into the preset correlation factor prediction model to obtain the target correlation factor. Based on the target correlation factor and the manufacturer's sales data, the terminal sales data missing for the specified period is obtained; Wherein, the manufacturer sales data and the terminal sales data are automobile sales data statistically analyzed based on different data sources, and the correlation factor prediction model is obtained based on the model training method of any one of claims 1 to 6.

8. The method according to claim 7, characterized in that, The manufacturer sales data includes total manufacturer sales data and manufacturer model sales data; the terminal sales data includes total terminal sales data and terminal model sales data; the target correlation factor includes total sales coefficient and model coefficient. The step of obtaining the terminal sales data missing for the specified period based on the target correlation factor and the manufacturer's sales data includes: Based on the total sales coefficient and the manufacturer's total sales data, the terminal total sales data missing for the specified period is obtained; Based on the vehicle model coefficient and the manufacturer's vehicle sales data, the terminal vehicle sales data missing for the specified period is obtained.

9. The method according to claim 8, characterized in that, The manufacturer vehicle sales data includes manufacturer first data and manufacturer second data corresponding to different vehicle models; the terminal vehicle sales data includes terminal first data and terminal second data corresponding to the different vehicle models; the vehicle model coefficient includes a first coefficient and a second coefficient corresponding to the different vehicle models; The step of obtaining the terminal vehicle sales data missing for the specified period based on the vehicle model coefficient and the manufacturer's vehicle sales data includes: Based on the first coefficient and the manufacturer's first data, the first estimated data is obtained; Based on the second coefficient and the manufacturer's second data, the second estimated data is obtained; The second remaining data is determined based on the total sales data of the terminals and the first estimated data; When the second remaining data is greater than or equal to the second estimated data, the first estimated data is used as the first data of the terminal, and the second remaining data is used as the second data of the terminal. When the second remaining data is less than the second estimated data, the first estimated data is adjusted according to the difference between the second remaining data and the second estimated data to obtain the terminal first data, and the second estimated data is used as the terminal second data.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.