Clothing sales prediction system based on big data analysis

By using Apache Spark for data integration and distributed computing in the process of clothing production and sales, combining clustering and association mining, and building a hybrid prediction model, the problems of low data processing efficiency and insufficient prediction accuracy in the existing technology are solved, and efficient clothing sales prediction is achieved.

CN120235643AActive Publication Date: 2025-07-01FUZHOU RONGZHIQUAN SOFTWARE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510297663.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate and process the data in the production and sales process of clothing, it is difficult to use Apache Spark for distributed processing, it is difficult to perform cluster analysis and association mining, and it is difficult to build a hybrid prediction model for clothing sales prediction.

Method used

By collecting data in the clothing production and sales process, using Apache Spark for distributed computing processing, combining clustering algorithms and Apriori algorithm for data mining, and building a hybrid prediction model for clothing sales prediction.

Benefits of technology

It has achieved efficient processing of large-scale clothing data, discovered potential relationships and influencing factors, and improved the accuracy and real-timeness of clothing sales forecasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235643A_ABST
    Figure CN120235643A_ABST
Patent Text Reader

Abstract

The invention discloses a garment sales prediction system based on big data analysis, relates to the technical field of data processing, and solves the problem that it is difficult to carry out distributed processing on collected and integrated data in garment production and sales processes by using Apache Spark; the clustering analysis is difficult to carry out on the Spark by utilizing a clustering algorithm; the correlation between the clothing picture data and the sales data is difficult to mine; an Apriori algorithm is difficult to mine the incidence relation; and a hybrid prediction model is difficult to construct for clothing sales prediction and real-time monitoring. According to the invention, by collecting, integrating and processing the data of the garment production and sales process, the distributed calculation processing module is utilized to efficiently process various garment data, the data mining model is utilized to deeply mine and analyze, and finally, the hybrid prediction model is constructed and the model parameters are updated in real time to carry out accurate sales prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and specifically relates to a clothing sales prediction system based on big data analysis. Background Art

[0002] In today's clothing industry, big data analysis technology plays a crucial role. Apache Spark is a fast and general computing engine that uses an in-memory computing model to store intermediate results in memory, thus significantly improving the processing speed. By leveraging Apache Spark, large-scale clothing datasets and image data can be efficiently processed. By applying clustering algorithms, correlation analysis methods, and the Apriori algorithm, etc., the potential information in multi-source clothing data is deeply mined, and the influence degrees in multiple aspects such as production, environment, sales season, and clothing styles are analyzed. Finally, a hybrid prediction model is constructed based on the results of data mining, so as to achieve accurate prediction of clothing sales and provide strong data support for the production and sales decisions of clothing enterprises.

[0003] The existing technologies have the following problems: it is difficult to perform distributed processing on the data in the clothing production and sales processes after collection, integration, and processing using Apache Spark; it is difficult to perform clustering analysis using clustering algorithms on Spark; it is difficult to mine the correlation between clothing image data and sales data; it is difficult to use the Apriori algorithm to mine association relationships; it is difficult to construct a hybrid prediction model for clothing sales prediction and real-time monitoring. Summary of the Invention

[0004] To solve the problems existing in the above-mentioned existing technologies, the present invention proposes a clothing sales prediction system based on big data analysis;

[0005] For this reason, the first aspect of the present invention provides a clothing sales prediction system based on big data analysis, including the following modules:

[0006] Data collection module: collect product data and production data in the clothing production process, and collect sales data, customer data, environmental data, and clothing image data in the clothing sales process;

[0007] Data processing module: integrate the collected product data, production data, sales data, customer data, and environmental data into a clothing dataset, and perform data cleaning and standardization on the integrated clothing dataset; perform filtering and denoising and image enhancement on the collected clothing image data;

[0008] Distributed computing processing module: import the clothing dataset and clothing image data into the Spark cluster through the use of Apache Spark for distributed computing processing;

[0009] Data mining module: perform clustering analysis on the clothing dataset by using clustering algorithms on Spark; perform correlation analysis on clothing image data and sales data; use the Apriori algorithm to mine association relationships based on the preprocessed clothing dataset, and analyze and calculate the production impact, environmental impact, sales season impact, and clothing style impact;

[0010] Prediction analysis module: build a hybrid prediction model for prediction analysis; calculate the relative deviation value between the actual and predicted sales data to update the parameters of the hybrid prediction model in real time.

[0011] Preferably, collect product data and production data during the clothing production process, and collect sales data, customer data, environmental data, and clothing image data during the clothing sales process, including the following steps:

[0012] Collect product data and production data during the clothing production process. Among them, the product data includes: cost, fabric composition, style design, and production process; the production data includes: production time, production quantity, production equipment startup rate, failure rate, and production efficiency;

[0013] Collect sales data, customer data, environmental data, and clothing image data during the clothing sales process. Among them, the sales data includes: sales time, sales quantity, sales amount, and profit; the customer data includes: customer quantity, stay time in different clothing areas, number of try-on times, clothing style selection, return rate, and customer satisfaction; the environmental data includes: temperature, humidity, wind speed, and light intensity.

[0014] Preferably, import the clothing dataset and clothing image data into the Spark cluster for distributed computing processing by using Apache Spark, including the following steps:

[0015] Build an Apache Spark cluster environment, and convert the preprocessed product data, production data, sales data, customer data, and environmental data into the DataFrame or Dataset format supported by Spark respectively; integrate TensorFlow onSpark for distributed processing of clothing image data;

[0016] Utilize the distributed storage and computing capabilities of Spark to disperse and store the clothing dataset and clothing image data on different nodes of the cluster;

[0017] Integrate the genetic algorithm scheduler into the task scheduling framework of Spark, and use the genetic algorithm scheduler to dynamically allocate tasks to different Spark nodes; by monitoring the execution of task scheduling, collect performance data including: task execution time, task waiting time, data transfer time, and resource utilization rate, and based on the performance data, through the iterative process of the genetic algorithm, obtain an optimized task scheduling and allocation strategy;

[0018] Use the distributed computing ability of Spark to parallel process various data in the clothing dataset; calculate statistical metrics of various data in the clothing dataset, including: maximum value, minimum value, mean value, median, and standard deviation;

[0019] According to the preprocessed clothing image data, use an image processing library to convert the clothing image data into a tensor representation; utilize the distributed computing framework of Spark to perform parallel feature extraction on the clothing image data; use a deep learning framework to build a convolutional neural network model to automatically extract edge features, texture features, clothing color features, and style features from the clothing image data; by training the convolutional neural network model using historical clothing image data, use the trained convolutional neural network model to automatically perform feature extraction on the real-time collected and preprocessed clothing image data and output the feature extraction results; splice the results of feature extraction to generate a feature vector.

[0020] Preferably, perform clustering analysis on the clothing dataset by using a clustering algorithm on Spark, including the following steps:

[0021] On Spark, use K-Means as the clustering algorithm; initially set the value of K through the elbow method; randomly select K initial centroid points; use the Euclidean distance method to calculate the distance from each data point to all centroid points, and assign each data point to the cluster represented by the nearest centroid point; recalculate the average value of all data points in each cluster as the new centroid point; by repeating the above steps until the centroid points no longer change or reach the maximum number of iterations;

[0022] According to the above steps, perform clustering analysis on the clothing dataset to obtain clusters of product data, production data, sales data, customer data, and environmental data; use the Spark MLlib library to calculate the silhouette coefficient of each data point, the average silhouette coefficient of the entire clothing dataset, and the average value of the silhouette coefficients of the data points in each cluster.

[0023] Preferably, perform correlation analysis on the clothing image data and sales data, including the following steps:

[0024] According to the convolutional neural network model on Spark, automatically extract features from the collected and preprocessed clothing image data and generate feature vectors. Use the Spark MLlib library for correlation analysis. The formula for calculating the correlation coefficient between the feature vectors and the sales data is:

[0025]

[0026] where G represents the correlation coefficient, X i is the i-th eigenvalue in the feature vector, Y i is the i-th sales data value, X 均 and Y 均 represent the mean of all eigenvalues in the feature vector and the mean of all sales data respectively.

[0027] Preferably, use the Apriori algorithm to mine association relationships based on the preprocessed clothing data set, and analyze and calculate the production impact degree, environmental impact degree, sales season impact degree, and clothing style impact degree, including the following steps:

[0028] Set the minimum support threshold to 0.1 and the minimum confidence threshold to 0.7; on the clothing data set, use the Apriori algorithm in the association rule mining algorithm to generate frequent item sets; the frequent item sets are filtered out with a support greater than or equal to the minimum support threshold;

[0029] Perform feature screening on the frequent item sets generated by the Apriori algorithm, and filter out the item sets containing production data features, sales season features, environmental data features, clothing style features, and sales data features;

[0030] Filter out association rules from the frequent item sets, and calculate the support, confidence, and lift for each generated association rule; filter out the association rules with a confidence greater than or equal to the set minimum confidence threshold according to the set minimum confidence threshold;

[0031] Calculate the production impact degree by multiplying the support of all production data features and sales data features by the lift of the production data features to the sales data features, and then accumulating the products;

[0032] Calculate the environmental impact degree by multiplying the support of all environmental data features and sales data features by the lift of the environmental data features to the sales data features, and then accumulating the products;

[0033] Calculate the sales season impact degree by multiplying the support of all sales season features and sales data features by the lift of the sales season features to the sales data features, and then accumulating the products;

[0034] The clothing style influence degree is obtained by calculating the product of the support degrees of all clothing style features and sales data features and multiplying it by the promotion degree of clothing style features on sales data features, and then cumulatively adding the products.

[0035] Preferably, a hybrid prediction model is constructed for prediction analysis, including the following steps:

[0036] Collect and preprocess the historical clothing dataset and historical clothing picture data, and through distributed computing processing and data mining analysis, obtain the statistical indicators, clustering analysis results, correlation analysis results, and association rule results of various data in the historical clothing dataset;

[0037] Construct a training dataset with the various analysis results obtained above. The training dataset includes: the statistical indicators of various data in the historical clothing dataset; the clustering of product data, production data, sales data, customer data, and environmental data, the silhouette coefficient of each data point, the average silhouette coefficient of the entire clothing dataset, and the average of the silhouette coefficients of data points in each cluster; the correlation coefficient between the historical clothing image feature vector and sales data; the production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree;

[0038] Use a long short-term memory network as a time series prediction model for analysis and prediction; arrange the training dataset in chronological order, and divide the data arranged in chronological order into multiple time windows, each time window containing the sales data within a time period; within each time window, label the sales volume, sales amount, and profit; input the labeled training dataset into the time series prediction model for training; input the real-time collected and preprocessed clothing dataset and clothing picture data into the time series prediction model, and output the predicted values of the sales data, which are: sales volume, sales amount, and profit;

[0039] Input the training dataset into the association prediction model for training; obtain the real-time training dataset through distributed computing processing and data mining analysis of the real-time collected and preprocessed clothing dataset and clothing picture data, and input the real-time training dataset into the association prediction model, and the output predicted values are: willingness to buy, purchase frequency, repurchase rate, sales volume, sales amount, and profit;

[0040] Label the clothing categories including: shirts, skirts, pants, and coats; input the labeled training dataset into the visual prediction model for training; input the real-time collected and preprocessed clothing dataset and clothing picture data into the visual prediction model, and output the categories of the predicted clothing picture data, and the predicted values are the probabilities of shirts, skirts, pants, and coats;

[0041] According to the predicted values respectively output by the time series prediction model, the correlation prediction model, and the visual prediction model after training is completed, and collect the actual values that are consistent with the time of the predicted values in real time;

[0042] Integrate the time series prediction model, the correlation prediction model, and the visual prediction model through the method of dynamic weighted fusion to obtain a hybrid prediction model; calculate the weights of the time series prediction model, the correlation prediction model, and the visual prediction model respectively according to the mean squared error of the data predicted by the time series prediction model, the correlation prediction model, and the visual prediction model; among them, the calculation formula of the mean squared error is: N is the total number of data in the collected historical clothing dataset and historical clothing pictures, x i is the actual value, y i is the predicted value;

[0043] By calculating the weight formula of the time series prediction model: Calculate the weight formula of the correlation prediction model: Calculate the weight formula of the visual prediction model:

[0044] Obtain the weight w t of the time series prediction model, the weight w g of the correlation prediction model, and the weight w s of the visual prediction model, where M1, M2, and M3 are the mean squared errors of the time series prediction model, the correlation prediction model, and the visual prediction model respectively;

[0045] Multiply the predicted values of the time series prediction model, the correlation prediction model, and the visual prediction model by the corresponding weights of the time series prediction model, the correlation prediction model, and the visual prediction model respectively, and perform weighted summation to obtain the predicted value of the hybrid prediction model;

[0046] According to the real-time collected and preprocessed clothing dataset and clothing picture data, input them into the training time series prediction model, the correlation prediction model, and the visual prediction model respectively, and input the predicted values of the three trained models into the hybrid prediction model to obtain the predicted value results of the hybrid prediction model including: the number of sales, the sales amount, and the profit.

[0047] Preferably, by analyzing and calculating the relative deviation value of the actual and predicted sales data, the parameters of the hybrid prediction model are updated in real time, including the following steps:

[0048] Establish a real-time monitoring mechanism, calculate the difference between the actual sales data and the predicted sales data, and then divide it by the predicted sales data to obtain the relative deviation value;

[0049] When the relative deviation value exceeds the standard deviation of historical sales data, a variant of stochastic gradient descent in the online learning algorithm is used to update the parameters of the hybrid prediction model in real time according to newly collected product data, production data, sales data, customer data, environmental data, and clothing picture data.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] By comprehensively collecting data in the clothing production and sales processes and using Apache Spark for distributed computing and processing, the present invention can efficiently process large-scale clothing data sets and picture data, greatly shortening the data processing time, improving the data processing efficiency, meeting the characteristics of large data volume and high processing requirements in the clothing industry, and providing strong support for fast and timely data mining and predictive analysis.

[0052] Through the data mining module, clustering algorithms, Apriori algorithms, etc. are used to deeply analyze the clothing data sets and picture data, and potential correlation relationships and influencing factors in the clothing data are mined, including: production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree, as well as the correlation between clothing data and picture data.

[0053] By integrating the time series prediction model, association prediction model, and visual prediction model using the method of dynamic weighted fusion, the present invention obtains a hybrid prediction model; through the hybrid prediction model for clothing sales prediction, and by updating the model parameters in real time, the clothing sales situation can be predicted more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0055] Figure 1 It is the system module diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0057] Please refer to Figure 1, an embodiment of the first aspect of the present invention provides a clothing sales prediction system based on big data analysis, including the following modules:

[0058] Data collection module: Collect product data and production data during the clothing production process, and collect sales data, customer data, environmental data, and clothing picture data during the clothing sales process;

[0059] Data processing module: Integrate the collected product data, production data, sales data, customer data, and environmental data into a clothing data set, and perform data cleaning and standardization on the integrated clothing data set; Filter and denoise the collected clothing picture data and perform image enhancement;

[0060] Distributed computing processing module: Import the clothing data set and clothing picture data into the Spark cluster using Apache Spark for distributed computing processing;

[0061] Data mining module: Perform clustering analysis on the clothing data set by using a clustering algorithm on Spark; Perform correlation analysis on the clothing picture data and sales data; Use the Apriori algorithm to mine association relationships based on the preprocessed clothing data set, and analyze and calculate the production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree;

[0062] Prediction analysis module: Build a hybrid prediction model for prediction analysis; Calculate the relative deviation value between the actual and predicted sales data to update the parameters of the hybrid prediction model in real time.

[0063] Specifically, during the clothing production process, product data and production data are tracked and recorded in real time, and sales data, customer data, environmental data, and clothing picture data during the clothing sales process are recorded. The collected product data, production data, sales data, customer data, and environmental data are integrated according to a unified format and standard to form a clothing data set. Data cleaning and standardization of the data in the clothing data set include: removing duplicate, incorrect, and incomplete data records, as well as unifying the time format and numerical range of the data, etc. Filtering and denoising are performed on the clothing picture data to remove noise points and interference in the images. Enhancement of the clarity and contrast of the images is carried out to facilitate subsequent analysis. The clothing data set and the clothing picture data are imported into the Apache Spark cluster, and the distributed computing power of Spark is utilized to quickly process and analyze large-scale data. Using the clustering algorithm on Spark, clustering analysis is performed on the clothing data set to identify different categories of clothing groups. Correlation analysis is carried out on the clothing picture data and the sales data to find the association between the picture features and the sales performance. The Apriori algorithm is used to mine the association rules in the clothing data set, analyze the influence of production, environment, sales season, and clothing style on sales, and calculate the production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree to provide a basis for decision-making. According to the analysis results of data mining, a hybrid prediction model is constructed to predict the clothing sales data. By analyzing and calculating the relative deviation value between the actual and predicted sales data, the parameters of the hybrid prediction model are dynamically adjusted.

[0064] In one embodiment of the present invention, collecting product data and production data during the clothing production process, and collecting sales data, customer data, environmental data, and clothing picture data during the clothing sales process includes the following steps:

[0065] Collecting product data and production data during the clothing production process, wherein the product data includes: cost, fabric composition, style design, and production process; the production data includes: production time, production quantity, production equipment startup rate, failure rate, and production efficiency;

[0066] Collecting sales data, customer data, environmental data, and clothing picture data during the clothing sales process, wherein the sales data includes: sales time, sales quantity, sales amount, and profit; the customer data includes: number of customers, staying time in different clothing areas, number of try-ons, clothing style selection, return rate, and customer satisfaction; the environmental data includes: temperature, humidity, wind speed, and light intensity.

[0067] Specifically, through a cost accounting tool, accurately record the raw material cost, labor cost, transportation cost and other related expenses of each piece of clothing. According to the information provided by the supplier or the laboratory test results, record the proportion of fabric components used in the clothing, such as cotton, polyester, wool, etc. The style design data is provided by the design team, covering design elements such as the clothing pattern, color, pattern, etc., and can be recorded in the form of a digital design platform or drawings. Record the production process data during the clothing processing, including: cutting, sewing, ironing, etc., as well as the technical standards and process requirements adopted. Use RFID radio frequency identification technology to record the specific start and end times of clothing production, and real-time statistics of the number of completed clothing. By installing sensors on production equipment, monitor the startup status, running duration and fault conditions of the equipment, and calculate the startup rate and failure rate. Combine the production time and production quantity to calculate the production quantity per unit time and evaluate the production efficiency. Record the sales time, sales quantity and sales amount through the POS system. Combine the sales amount and cost data to calculate the profit of each piece of clothing, and real-time update and record the quantity of various types of clothing in the inventory. Statistically count the number of customers who purchase clothing, and by installing surveillance cameras or using intelligent sensing devices in the sales venue, record the staying time, try-on times, etc. of customers in different clothing areas. Through sales records or customer feedback, collect customer preference information for clothing styles. Record customer return requests, and collect customer satisfaction information through questionnaires or online evaluations. Install temperature and humidity sensors in the sales venue to monitor and record the temperature and humidity of the environment in real time; use anemometer or weather station data to record the wind speed in the sales venue; install light sensors to monitor the lighting level in the sales venue to ensure good display effects. The clothing image data including details such as the style, color, and material of the displayed clothing is collected by a photographer or the design team.

[0068] In one embodiment of the present invention, by using Apache Spark, import the clothing dataset and clothing picture data into the Spark cluster for distributed computing processing, including the following steps:

[0069] Build an Apache Spark cluster environment, and convert the preprocessed product data, production data, sales data, customer data and environmental data into DataFrame or Dataset formats supported by Spark respectively; integrate TensorFlow onSpark for distributed processing of clothing picture data;

[0070] Utilize the distributed storage and computing capabilities of Spark to disperse and store the clothing dataset and clothing picture data on different nodes of the cluster;

[0071] Integrate the genetic algorithm scheduler into the task scheduling framework of Spark, and use the genetic algorithm scheduler to dynamically allocate tasks to different Spark nodes; by monitoring the execution of task scheduling, collect performance data including: task execution time, task waiting time, data transmission time, and resource utilization rate, and based on the performance data, through the iterative process of the genetic algorithm, obtain an optimized task scheduling and allocation strategy;

[0072] Use the distributed computing ability of Spark to parallel process various data in the clothing dataset; calculate statistical metrics of various data in the clothing dataset, including: maximum value, minimum value, mean, median, and standard deviation;

[0073] According to the preprocessed clothing image data, use an image processing library to convert the clothing image data into a tensor representation; utilize the distributed computing framework of Spark to perform parallel feature extraction on the clothing image data; use a deep learning framework to build a convolutional neural network model to automatically extract edge features, texture features, clothing color features, and style features from the clothing image data; by training the convolutional neural network model using historical clothing image data, use the trained convolutional neural network model to automatically perform feature extraction on the real-time collected and preprocessed clothing image data and output the feature extraction results; splice the results of feature extraction to generate a feature vector.

[0074] Specifically, download and configure Spark from the official Apache Spark website. Set the master node address, serialization method, memory configuration, etc. of Spark in spark-defaults.conf. Set environment variables such as JAVA_HOME and HADOOP_HOME in spark-env.sh. Add the information of Worker nodes in the workers file. Use the scp command to distribute the installation directory of Spark to other nodes in the cluster and add the Spark environment variables on all nodes. Preprocess product data, production data, sales data, customer data, and environmental data, including data cleaning, format conversion, and missing value handling, etc. Convert the preprocessed data into the DataFrame or Dataset format supported by Spark for subsequent distributed computing processing. Install TensorFlow on each node of the Spark cluster and configure the environment of TensorFlow onSpark so that it can run on the Spark cluster. Utilize the distributed computing ability of TensorFlow on Spark to process clothing image data, such as image enhancement, image scaling, etc. Disperse and store the clothing dataset and clothing image data on different nodes of the Spark cluster to achieve distributed data storage. Use the distributed computing ability of Spark to process the data stored in the cluster in parallel to improve the data processing efficiency. Integrate the genetic algorithm scheduler into the task scheduling framework of Spark. Use the genetic algorithm scheduler to dynamically allocate tasks to different Spark nodes to optimize the task execution efficiency. Monitor the execution of task scheduling and collect performance data, including task execution time, task waiting time, data transfer time, and resource utilization, etc. According to the performance data, through the iterative process of the genetic algorithm, obtain the optimized task scheduling and allocation strategy. Use the distributed computing ability of Spark to process various data in the clothing dataset in parallel. Calculate the statistical metrics of various data in the clothing dataset, including maximum value, minimum value, mean, median, and standard deviation, etc. According to the preprocessed clothing image data, use the image processing library TensorFlow to convert the clothing image data into a tensor representation, and utilize the distributed computing framework of Spark to perform parallel feature extraction on the clothing image data. Use the deep learning framework to build a convolutional neural network model to automatically extract edge features, texture features, clothing color features, and style features in the clothing image data. Use the historical clothing image data to train the convolutional neural network model. Use the trained convolutional neural network model to automatically perform feature extraction on the real-time collected and preprocessed clothing image data and output the feature extraction results. Concatenate the results of feature extraction to generate a feature vector.Among them, the specifically extracted feature parameters include: in color features, the proportion of pixels of different colors and the color distribution in the picture are described by a color histogram, as well as the visual weight of different colors in the whole picture, and the saturation and brightness of the clothing color are extracted. In texture features, the determined texture types include: cotton texture, silk texture, denim texture, etc. The thickness of each texture type is extracted, which represents the size and graininess of the texture, and the orientation and distribution law of each texture type in the clothing picture, including: the direction of stripe texture is horizontal, vertical or oblique, the size and arrangement of checkered patterns, etc. The surface area of the outer contour shape of the clothing main body is extracted, and the internal shape includes the surface area of the shapes of components such as patterns, pockets, collars, cuffs, etc. on the clothing. In edge features, the edge sharpness and edge smoothness are extracted.

[0075] In one embodiment of the present invention, by using a clustering algorithm on Spark, a clustering analysis is performed on the clothing data set, including the following steps:

[0076] On Spark, use K-Means as the clustering algorithm; initially set the value of K by the elbow method; randomly select K initial centroid points; calculate the distance from each data point to all centroid points using the Euclidean distance method, and assign each data point to the cluster represented by the nearest centroid point; recalculate the average value of all data points in each cluster as the new centroid point; by repeating the above steps until the centroid points no longer change or reach the maximum number of iterations;

[0077] According to the above steps, a clustering analysis is performed on the clothing data set to obtain the clustering of product data, the clustering of production data, the clustering of sales data, the clustering of customer data, and the clustering of environmental data; use the Spark MLlib library to calculate the silhouette coefficient of each data point, the average silhouette coefficient of the entire clothing data set, and the average value of the silhouette coefficients of the data points in each cluster.

[0078] Specifically, ensure that there is a configured Spark cluster, including: a Master node and multiple Worker nodes. Upload the clothing dataset to HDFS or other distributed storage systems so that Spark can access and process it efficiently. Add a dependency on MLlib, which is Spark's machine learning library, to the Spark project. First, by plotting the SSE (Sum of Squared Errors) graph for different K values, observe the elbow position to determine a suitable K value. The elbow position is usually the point where the SSE starts to decline sharply and then becomes flat, and the K value corresponding to this point is considered the optimal number of clusters. Randomly select K data points from the clothing dataset as the initial centroid points. For each data point in the clothing dataset, calculate the Euclidean distance from each data point to all K centroid points; assign each data point to the cluster represented by the centroid point closest to it. For each cluster, calculate the average value of all the data points within it as the new centroid point. Repeat the above steps until the positions of the centroid points no longer change significantly or reach the preset maximum number of iterations. According to the clustering results, divide the clothing dataset into clusters of product data, production data, sales data, customer data, and environmental data. Use the silhouette coefficient in the Spark MLlib library to evaluate the quality of the clustering results. The silhouette coefficient measures the similarity of a data point to other points within its assigned cluster and the dissimilarity to points in other clusters. Calculate the silhouette coefficient for each data point, calculate the average silhouette coefficient for the entire clothing dataset, and calculate the average of the silhouette coefficients of the data points in each cluster.

[0079] In one embodiment of the present invention, the correlation analysis of clothing picture data and sales data includes the following steps:

[0080] According to the convolutional neural network model on Spark, automatically extract features from the collected and preprocessed clothing picture data and generate feature vectors. Use the Spark MLlib library for correlation analysis. The formula for calculating the correlation coefficient between the feature vectors and the sales data is:

[0081]

[0082] where G represents the correlation coefficient, X i is the i-th eigenvalue in the feature vector, Y i is the i-th sales data value, X 均 and Y 均 represent the mean of all eigenvalues in the feature vector and the mean of all sales data, respectively.

[0083] Specifically, a convolutional neural network model on Spark is used to extract features from the preprocessed clothing image data. The convolutional neural network model automatically learns and extracts features in the image through structures such as convolutional layers and pooling layers, including: color, edges, texture, shape, etc. After feature extraction is completed, the image data is converted into feature vectors for subsequent correlation analysis. Using the correlation analysis tools provided by the Spark MLlib library, correlation analysis is performed on the feature vectors and sales data. By substituting the extracted clothing image features and the values in the sales data into the correlation coefficient formula between the feature vectors and the sales data, the correlation coefficient is obtained; the correlation coefficient is used to measure the linear correlation degree between the feature vectors and the sales data. The closer the value of the correlation coefficient is to 1 or -1, the stronger the correlation between the two; the closer the value of the correlation coefficient is to 0, the weaker the correlation between the two.

[0084] Based on the calculated correlation coefficient, it can be determined which clothing image features have a significant impact on the sales data.

[0085] In one embodiment of the present invention, the Apriori algorithm is used to mine association relationships according to the preprocessed clothing data set, and the production impact degree, environmental impact degree, sales season impact degree, and clothing style impact degree are analyzed and calculated, including the following steps:

[0086] Set the minimum support threshold to 0.1 and the minimum confidence threshold to 0.7; on the clothing data set, use the Apriori algorithm in the association rule mining algorithm to generate frequent item sets; the frequent item sets are obtained by screening with a support greater than or equal to the minimum support threshold;

[0087] Feature screening is performed on the frequent item sets generated by the Apriori algorithm, and item sets containing production data features, sales season features, environmental data features, clothing style features, and sales data features are screened out;

[0088] Association rules are screened out from the frequent item sets, and the support, confidence, and lift are calculated for each generated association rule; association rules with a confidence greater than or equal to the set minimum confidence threshold are screened out according to the set minimum confidence threshold;

[0089] By calculating the product of the support of all production data features and sales data features multiplied by the lift of the production data features to the sales data features, and then cumulatively adding the products, the production impact degree is obtained;

[0090] By calculating the product of the support of all environmental data features and sales data features multiplied by the lift of the environmental data features to the sales data features, and then cumulatively adding the products, the environmental impact degree is obtained;

[0091] The sales season impact is obtained by calculating the product of the support degrees of all sales season features and sales data features, multiplying it by the promotion degree of the sales season features on the sales data features, and then cumulatively adding up the products.

[0092] The clothing style impact is obtained by calculating the product of the support degrees of all clothing style features and sales data features, multiplying it by the promotion degree of the clothing style features on the sales data features, and then cumulatively adding up the products.

[0093] Specifically, set the minimum support threshold to 0.1. This threshold is used to screen frequent item sets, and only item sets with a support greater than or equal to this threshold will be retained. Set the minimum confidence threshold to 0.7. This threshold is used to screen association rules, and only rules with a confidence greater than or equal to this threshold will be regarded as valid rules. Among them, the minimum support threshold and the minimum confidence threshold are determined by adjusting through experiments or actual situations. An initial value can be set, and then the algorithm is run to observe the quantity and quality of the frequent item sets and association rules in the results. If there are too many frequent item sets, the support threshold can be appropriately increased; if there are too few association rules, the confidence threshold can be appropriately decreased. Apply the Apriori algorithm to the preprocessed clothing dataset to generate frequent item sets. Frequent item sets are generated iteratively, starting from frequent 1-item sets and gradually generating higher-order frequent item sets. Screen out the item sets containing production data features, sales season features, environmental data features, clothing style features, and sales data features from the generated frequent item sets. Among them, the feature parameters in the production data features, environmental data features, clothing style features, and sales data features are the collected production data, environmental data, clothing styles, and sales data respectively. The sales season feature divides the clothing sales season into spring, summer, autumn, and winter according to the time nodes of the four seasons of the year. Generate association rules from the frequent item sets and calculate the support, confidence, and lift of each association rule. Screen out the association rules with a confidence greater than or equal to the set minimum confidence threshold. Calculate the product of the support and lift of all association rules containing production data features and sales data features, and accumulate and sum these products to obtain the production influence degree. This indicator reflects the degree of influence of production factors on sales data. Similarly, calculate the product of the support and lift of all association rules containing environmental data features and sales data features, and accumulate and sum these products to obtain the environmental influence degree. This indicator reflects the degree of influence of environmental factors on sales data. Calculate the product of the support and lift of all association rules containing sales season features and sales data features, and accumulate and sum these products to obtain the sales season influence degree. This indicator reflects the degree of influence of the sales season on sales data. Calculate the product of the support and lift of all association rules containing clothing style features and sales data features, and accumulate and sum these products to obtain the clothing style influence degree. This indicator reflects the degree of influence of clothing styles on sales data. Explain the degree of influence of different factors on sales data based on the calculated production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree.

[0094] In one embodiment of the present invention, a hybrid prediction model is constructed for prediction analysis, including the following steps:

[0095] The historical clothing dataset and historical clothing picture data after collection and preprocessing are processed through distributed computing and data mining analysis to obtain statistical indicators, clustering analysis results, correlation analysis results, and association rule results of various data in the historical clothing dataset;

[0096] Construct a training dataset from the various analysis results obtained above. The training dataset includes: statistical indicators of various data in the historical clothing dataset; clustering of product data, production data, sales data, customer data, and environmental data, silhouette coefficients of each data point, average silhouette coefficient of the entire clothing dataset, and average of silhouette coefficients of data points in each cluster; correlation coefficient between historical clothing image feature vectors and sales data; production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree;

[0097] Use a long short-term memory network as a time series prediction model for analysis and prediction; arrange the training dataset in chronological order and divide the data arranged in chronological order into multiple time windows, each time window containing sales data within a time period; label the sales quantity, sales amount, and profit within each time window; input the labeled training dataset into the time series prediction model for training; input the real-time collected and preprocessed clothing dataset and clothing picture data into the time series prediction model, and output the predicted values of sales data, which are: sales quantity, sales amount, and profit;

[0098] Input the training dataset into the association prediction model for training; obtain the real-time training dataset by processing the real-time collected and preprocessed clothing dataset and clothing picture data through distributed computing and data mining analysis, and input the real-time training dataset into the association prediction model, and the output predicted values are: willingness to buy, purchase frequency, repurchase rate, sales quantity, sales amount, and profit;

[0099] Label the clothing categories including: shirts, skirts, pants, and coats; input the labeled training dataset into the visual prediction model for training; input the real-time collected and preprocessed clothing dataset and clothing picture data into the visual prediction model, and output the category of the predicted clothing picture data, and the predicted values are the probabilities of shirts, skirts, pants, and coats;

[0100] According to the predicted values respectively output by the trained time series prediction model, association prediction model, and visual prediction model, and collect the actual values consistent with the time of the predicted values in real time;

[0101] Integrate the time series prediction model, the correlation prediction model, and the visual prediction model through the method of dynamic weighted fusion to obtain a hybrid prediction model; calculate the weights of the time series prediction model, the correlation prediction model, and the visual prediction model respectively according to the mean square error of the data predicted by the time series prediction model, the correlation prediction model, and the visual prediction model; among them, the calculation formula of the mean square error is: N is the total number of data in the collected historical clothing dataset and historical clothing pictures, x i is the actual value, y i is the predicted value;

[0102] Calculate the weight formula of the time series prediction model: Calculate the weight formula of the correlation prediction model: Calculate the weight formula of the visual prediction model:

[0103] Obtain the weight w of the time series prediction model t 、the weight w of the correlation prediction model g and the weight w of the visual prediction model s , where M1, M2, and M3 are the mean square errors of the time series prediction model, the correlation prediction model, and the visual prediction model respectively;

[0104] Multiply the predicted values of the time series prediction model, the correlation prediction model, and the visual prediction model by the corresponding weights of the time series prediction model, the correlation prediction model, and the visual prediction model respectively, and perform weighted summation to obtain the predicted value of the hybrid prediction model;

[0105] According to the real-time collected and preprocessed clothing dataset and clothing picture data, input them into the trained time series prediction model, correlation prediction model, and visual prediction model respectively, and input the predicted values of the three trained models into the hybrid prediction model to obtain the predicted value results of the hybrid prediction model, including: sales volume, sales amount, and profit.

[0106] Specifically, information is collected from the historical clothing dataset and historical clothing picture data, including but not limited to product data, production data, sales data, customer data, and environmental data. Preprocessing steps such as data cleaning, denoising, and normalization are performed on the data to ensure data quality. Various statistical metrics in the historical clothing dataset are calculated, such as mean, standard deviation, maximum value, minimum value, etc. Clustering algorithms are used to cluster product data, production data, sales data, customer data, and environmental data, and the silhouette coefficient of each data point and the average silhouette coefficient of the entire dataset are calculated. The correlation coefficient between the historical clothing image feature vectors and sales data is analyzed. Association rules in the historical data are mined to obtain the production influence degree, environmental influence degree, sales season influence degree, and clothing style influence degree. The long short-term memory network is used as a time series prediction model. The training dataset is arranged in chronological order and divided into multiple time windows. The sales quantity, sales amount, and profit are labeled within each time window. The labeled training dataset is input into the long short-term memory network model for training. Distributed computing and data mining techniques are used to process real-time data to obtain a real-time training dataset, and the real-time training dataset is input into the association prediction model for training. The labeled clothing categories include but are not limited to shirts, skirts, pants, and coats. The labeled training dataset is input into the convolutional neural network of the visual prediction model for training. According to the actual values in the historical dataset and the predicted values of each model, the mean squared error is calculated. The weights of the time series prediction model, association prediction model, and visual prediction model are calculated based on the mean squared error. The predicted values of each model are multiplied by the corresponding weights to obtain the predicted value of the hybrid prediction model. The preprocessed clothing dataset and clothing picture data collected in real time are input into each model and the predicted values of each model are output. The predicted values of each model are input into the hybrid prediction model to obtain the final prediction results, including but not limited to sales quantity, sales amount, and profit.

[0107] In one embodiment of the present invention, by analyzing and calculating the relative deviation value between the actual and predicted sales data, the parameters of the hybrid prediction model are updated in real time, including the following steps:

[0108] A real-time monitoring mechanism is established. The relative deviation value is obtained by calculating the difference between the actual sales data and the predicted sales data and then dividing it by the predicted sales data.

[0109] When the relative deviation value exceeds the standard deviation of the historical sales data, a variant of stochastic gradient descent in the online learning algorithm is used to update the parameters of the hybrid prediction model in real time according to the newly collected product data, production data, sales data, customer data, environmental data, and clothing picture data.

[0110] Specifically, by calculating the difference between the actual sales data and the predicted sales data, and then dividing it by the predicted sales data, a relative deviation value is obtained, which reflects the degree of difference between the predicted sales data and the actual sales data. A threshold is set, and this threshold can be determined according to the standard deviation of historical sales data. The standard deviation is an important indicator to measure the volatility of data. When the relative deviation value exceeds this threshold, it means that the hybrid prediction model needs to be updated. By monitoring the relative deviation value in real time, when it is found that it exceeds the threshold, the model update mechanism is triggered. In this embodiment, a variant of stochastic gradient descent in the online learning algorithm is selected to update the parameters of the hybrid prediction model in real time. Stochastic gradient descent is an optimization algorithm suitable for large-scale datasets and online learning scenarios. When the relative deviation value exceeds the threshold, new product data, production data, sales data, customer data, environmental data, and clothing picture data are collected. These data are integrated into the model to provide rich information to support the update of the model. Using the variant algorithm of stochastic gradient descent, the parameters of the hybrid prediction model are updated in real time according to the newly collected data, and the prediction error is reduced by continuously adjusting the model parameters.

[0111] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A clothing sales forecasting system based on big data analysis, characterized in that: Includes the following modules: Data collection module: collects product data and production data in the clothing production process, and collects sales data, customer data, environmental data and clothing image data in the clothing sales process; Data processing module: Integrate the collected product data, production data, sales data, customer data and environmental data into a clothing data set, and perform data cleaning and standardization on the integrated clothing data set; Perform filtering, denoising and image enhancement on the collected clothing image data; Distributed computing processing module: By using Apache Spark, the clothing data set and clothing image data are imported into the Spark cluster for distributed computing processing; Data mining module: cluster analysis is performed on clothing data sets by using clustering algorithms on Spark; correlation analysis is performed on clothing image data and sales data; the Apriori algorithm is used to mine association relationships based on the preprocessed clothing data sets, and the production impact, environmental impact, sales season impact, and clothing style impact are analyzed and calculated; Prediction and analysis module: Build a hybrid prediction model for prediction and analysis; update the hybrid prediction model parameters in real time by calculating the relative deviation between actual and predicted sales data.

2. The clothing sales forecasting system based on big data analysis according to claim 1 is characterized in that: Collecting product data and production data in the clothing production process, and collecting sales data, customer data, environmental data and clothing image data in the clothing sales process, including the following steps: Collect product data and production data during the clothing production process, where product data includes: cost, fabric composition, style design and production process; production data includes: production time, production quantity, production equipment start-up rate, failure rate and production efficiency; Collect sales data, customer data, environmental data and clothing image data during the clothing sales process. Sales data includes: sales time, sales volume, sales revenue and profit; customer data includes: number of customers, stay time in different clothing areas, number of tries, clothing style selection, return rate and customer satisfaction; environmental data includes: temperature, humidity, wind speed and light intensity.

3. The clothing sales forecasting system based on big data analysis according to claim 1 is characterized in that: By using Apache Spark, the clothing dataset and clothing image data are imported into the Spark cluster for distributed computing processing, including the following steps: Build an Apache Spark cluster environment to convert pre-processed product data, production data, sales data, customer data, and environmental data into DataFrame or Dataset formats supported by Spark; integrate TensorFlow on Spark to process distributed clothing image data; Using Spark's distributed storage and computing capabilities, the clothing dataset and clothing image data are stored in different nodes of the cluster; Integrate the genetic algorithm scheduler into Spark's task scheduling framework, and use the genetic algorithm scheduler to dynamically assign tasks to different Spark nodes; monitor the execution of task scheduling, collect performance data including task execution time, task waiting time, data transmission time and resource utilization, and obtain the optimized task scheduling allocation strategy based on the performance data through the iterative process of the genetic algorithm; Use Spark's distributed computing capabilities to parallelize the various data in the clothing dataset; calculate the statistical indicators of various data in the clothing dataset, including: maximum value, minimum value, mean, median, and standard deviation; According to the preprocessed clothing image data, the image processing library is used to convert the clothing image data into tensor representation; the Spark distributed computing framework is used to perform parallel feature extraction on the clothing image data; a convolutional neural network model is built using a deep learning framework to automatically extract edge features, texture features, clothing color features and style features from the clothing image data; the convolutional neural network model is trained by using historical clothing image data, and the trained convolutional neural network model is used to automatically extract features from the real-time collected and preprocessed clothing image data and output the feature extraction results; the feature extraction results are spliced ​​to generate feature vectors.

4. The clothing sales forecasting system based on big data analysis according to claim 3 is characterized in that: By using clustering algorithms on Spark, clustering analysis is performed on the clothing dataset, including the following steps: On Spark, K-Means is used as the clustering algorithm. The K value is initially set by the elbow rule. K initial centroids are randomly selected. The distance from each data point to all centroids is calculated using the Euclidean distance method, and each data point is assigned to the cluster represented by the nearest centroid. The average value of all data points in each cluster is recalculated as the new centroid. The above steps are repeated until the centroid no longer changes or the maximum number of iterations is reached. According to the above steps, cluster analysis is performed on the clothing dataset to obtain clustering of product data, clustering of production data, clustering of sales data, clustering of customer data, and clustering of environmental data; the Spark MLli b library is used to calculate the silhouette coefficient of each data point, the average silhouette coefficient of the entire clothing dataset, and the average silhouette coefficient of the data points in each cluster.

5. The clothing sales forecasting system based on big data analysis according to claim 4 is characterized in that: The correlation analysis of clothing image data and sales data includes the following steps: According to the convolutional neural network model on Spark, the collected and preprocessed clothing image data is automatically extracted and feature vectors are generated. The Spark MLlib library is used for correlation analysis. The correlation coefficient formula between the feature vector and the sales data is calculated as follows: Among them, G represents the correlation coefficient, X i is the i-th eigenvalue in the eigenvector, Y i is the i-th sales data value, X 均 and Y 均 They represent the mean of all eigenvalues ​​in the eigenvector and the mean of all sales data respectively.

6. The clothing sales forecasting system based on big data analysis according to claim 1 is characterized in that: The Apriori algorithm is used to mine the association relationship based on the preprocessed clothing data set, and the production impact, environmental impact, sales season impact and clothing style impact are analyzed and calculated, including the following steps: The minimum support threshold is set to 0.1 and the minimum confidence threshold is set to 0.7; on the clothing data set, the Apriori algorithm in the association rule mining algorithm is used to generate frequent item sets; the frequent item sets are obtained by screening with support greater than or equal to the minimum support threshold; Perform feature screening based on the frequent item sets generated by the Apriori algorithm to screen out item sets containing production data features, sales season features, environmental data features, clothing style features, and sales data features; Filter out association rules from frequent item sets, and calculate support, confidence, and lift for each generated association rule; filter out association rules whose confidence is greater than or equal to the minimum confidence threshold according to the set minimum confidence threshold; The production impact is obtained by calculating the support of all production data features and sales data features multiplied by the product of the improvement of production data features on sales data features, and then adding the products together; The environmental impact is obtained by calculating the support of all environmental data features and sales data features and multiplying the product of the environmental data features' enhancement of the sales data features, and then adding the products together. The sales season influence is obtained by calculating the product of the support of all sales season features and sales data features and the improvement of sales season features on sales data features, and then adding the products. The influence of clothing styles is obtained by calculating the product of the support of all clothing style features and sales data features and multiplying the improvement of clothing style features on sales data features, and then adding the products.

7. The clothing sales forecasting system based on big data analysis according to claim 1 is characterized in that: Building a hybrid prediction model for predictive analysis includes the following steps: Collect and preprocess historical clothing data sets and historical clothing image data, and obtain statistical indicators, cluster analysis results, correlation analysis results, and association rule results of various data in the historical clothing data sets through distributed computing processing and data mining analysis; The various analysis results obtained above are used to construct a training data set, which includes: statistical indicators of various data in the historical clothing data set; clustering of product data, clustering of production data, clustering of sales data, clustering of customer data and clustering of environmental data, as well as the silhouette coefficient of each data point, the average silhouette coefficient of the entire clothing data set and the average of the silhouette coefficients of the data points in each cluster; the correlation coefficient between the historical clothing image feature vector and the sales data; the production influence, environmental influence, sales season influence and clothing style influence; Use the long short-term memory network as a time series prediction model for analysis and prediction; arrange the training data set in chronological order, and divide the chronologically arranged data into multiple time windows, each of which contains sales data within a time period; in each time window, annotate the sales quantity, sales amount, and profit; input the labeled training data set into the time series prediction model for training; input the clothing data set and clothing image data collected and preprocessed in real time into the time series prediction model, and output the predicted sales data value, the predicted value is: sales quantity, sales amount, and profit; The training data set is input into the association prediction model for training; the clothing data set and clothing image data collected and preprocessed in real time are processed by distributed computing and data mining analysis to obtain a real-time training data set, and the real-time training data set is input into the association prediction model, and the output prediction values ​​are: purchase intention, purchase frequency, repurchase rate, sales quantity, sales amount and profit; The labeled clothing categories include shirts, skirts, pants, and coats; the labeled training data set is input into the visual prediction model for training; the clothing data set and clothing image data collected and preprocessed in real time are input into the visual prediction model, and the predicted category of the clothing image data is output, and the predicted value is the probability of a shirt, the probability of a skirt, the probability of pants, and the probability of a coat; According to the prediction values ​​output by the time series prediction model, association prediction model and visual prediction model after training, the actual values ​​consistent with the prediction value time are collected in real time; The time series prediction model, the association prediction model and the visual prediction model are integrated by the dynamic weighted fusion method to obtain a hybrid prediction model; according to the mean square error of the prediction data of the time series prediction model, the association prediction model and the visual prediction model, the weights of the time series prediction model, the association prediction model and the visual prediction model are calculated respectively; wherein, the calculation formula of the mean square error is: N is the total number of historical costume data sets and historical costume pictures collected, x i is the actual value, y i is the predicted value; By calculating the weight formula of the time series prediction model: The weight formula for calculating the association prediction model is: The weight formula for calculating the visual prediction model is: Get the weight w of the time series prediction model t , the weight w of the association prediction model g and the weights w of the visual prediction model s , where M1, M2 and M3 are the mean square errors of the time series prediction model, the association prediction model and the visual prediction model respectively; The prediction values ​​of the time series prediction model, the association prediction model and the visual prediction model are multiplied by the corresponding weights of the time series prediction model, the association prediction model and the visual prediction model, respectively, and the weighted sum is obtained to obtain the prediction value of the hybrid prediction model; According to the clothing data set and clothing image data collected and preprocessed in real time, the time series prediction model, the association prediction model and the visual prediction model are input into the training respectively, and the prediction values ​​of the three models obtained by training are input into the hybrid prediction model. The prediction value results of the hybrid prediction model include: sales quantity, sales amount and profit.

8. The clothing sales forecasting system based on big data analysis according to claim 1 is characterized in that: By calculating the relative deviation between the actual and predicted sales data, the hybrid forecasting model parameters are updated in real time, including the following steps: Establish a real-time monitoring mechanism to calculate the difference between actual sales data and predicted sales data, and then divide it by the predicted sales data to obtain the relative deviation value; When the relative deviation value exceeds the standard deviation of historical sales data, a variant of stochastic gradient descent in the online learning algorithm is used to update the hybrid prediction model parameters in real time based on newly collected product data, production data, sales data, customer data, environmental data, and clothing image data.

Citation Information

Patent Citations

  • Enterprise clothing productivity prediction method and system

    CN118313871A

  • Sales and traffic data analysis

    US20200334696A1