Wine consumption prediction method combining curvelet transform and convolutional neural network
By combining the methods of curved wave transformation and convolutional neural network, multi-scale feature decomposition and local feature extraction are carried out, and the accuracy and stability problems in alcohol consumption prediction are solved, achieving more accurate consumer behavior analysis and market prediction.
Patent Information
- Application Number
- CN202510886233.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing alcohol consumption prediction methods have shortcomings in accuracy and stability, and have failed to fully tap potential information in consumer behavior data, and lack effective decomposition and correlation analysis of data characteristics.
Combining curve wave transformation and convolutional neural network, multi-scale feature decomposition and feature extraction are carried out by constructing a time-feature data set, local feature extraction and pooling operations are used for convolutional neural network, and alcohol consumption prediction is finally carried out through linear correction activation.
It significantly improves the accuracy and stability of alcohol consumption forecasts, can capture consumers' behavior patterns more comprehensively, adapt to complex and changeable market environments, reduce noise interference, and improve the intelligence level of the model.
Smart Images

Figure CN120387846A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of specific computing models, and particularly relates to a method for predicting liquor consumption by combining curvelet transform and convolutional neural network. Background Art
[0002] With the rapid development of information technology, consumption prediction, as a market operation model, has increasingly attracted the attention of various enterprises. Especially in the liquor sales field, accurately predicting consumers' purchase possibilities is of great significance for improving marketing efficiency, optimizing inventory management, and increasing enterprise revenue. By analyzing consumers' historical data, including basic profile information, daily behavior trajectories, and purchase results, etc., it can help merchants better understand consumers' behavior patterns and formulate precise marketing strategies accordingly.
[0003] Currently, the prediction of liquor consumers' purchase possibilities mainly relies on two methods, namely prediction based on simple mathematical models and prediction based on large artificial intelligence models. Among them, the prediction method based on simple mathematical models usually uses probability models, transformation models, etc. to analyze the probability distribution and trends of data. Although this method is easy to operate, the degree of data mining it processes is relatively low, and it fails to fully utilize the correlation characteristics and hidden information between consumers' behavior data, resulting in low prediction accuracy and only providing rough guidance in the process of precision marketing.
[0004] In addition, the prediction method based on large artificial intelligence models uses large-scale data and powerful computing capabilities to train "large-parameter" models, which have high generalization. However, such models lack in-depth analysis of the rationality of liquor consumption data and research on feature correlations in specific applications, especially perform poorly when dealing with data with unclear features, making the prediction stability of the models poor and difficult to convince enterprises in actual operations.
[0005] Although the above methods have their respective advantages, there are still significant deficiencies. Therefore, there is an urgent need for a new method to solve the problems of insufficient prediction accuracy and poor model stability in the prior art. Summary of the Invention
[0006] The present invention provides a method for predicting liquor consumption by combining curvelet transform and convolutional neural network to solve the problems that the potential information in consumers' behavior data fails to be fully mined, resulting in inaccurate prediction results, and the model performs unstably in different application scenarios due to the lack of effective decomposition and correlation analysis of data features.
[0007] The technical solution adopted by the present invention is as follows: A method for predicting liquor consumption by combining curvelet transform and convolutional neural network, comprising: Based on the user basic profiles, daily behavior trajectories, and purchase information results of alcohol consumption users, a time-feature dataset is obtained; According to the time-characteristic data set, a plurality of curvelet coefficients are obtained by curvelet transformation in combination with the curvelet characteristic frequency, wherein the plurality of curvelet coefficients include a low-frequency curvelet coefficient, a medium-frequency curvelet coefficient, a high-frequency curvelet coefficient, and an ultra-high-frequency curvelet coefficient; According to the low-frequency curvelet coefficient, the intermediate-frequency curvelet coefficient, and the high-frequency curvelet coefficient, a plurality of feature information is obtained by feature decomposition; Based on the multiple feature information, a convolutional neural network is used to obtain multiple convolution operation results, a pooling operation result is obtained through pooling operation, and a wine consumption prediction is obtained through linear correction activation.
[0008] The alcohol consumption prediction method combining curvelet transform and convolutional neural network described in the present invention also includes the following additional technical features: Based on the user basic profiles, daily behavior trajectories, and purchase information results of alcohol consumption users, a time-feature dataset is obtained, specifically: Based on the basic user profiles, daily behavior trajectories, and purchase information of alcohol consumers, format and transform them to obtain time series values in a one-to-one correspondence; A time-feature data set is obtained according to the time series values.
[0009] The basic user profile, daily behavior trajectory, and purchase information results of alcohol consumption users are as follows: The basic user profile includes at least one of the consumer's age, gender, date of birth, and registration time; Daily behavior traces include at least one of the following: time logged into the platform, time spent browsing products, time initiating payment, number of order cancellations, number of product collections, and number of activity participations; The purchase result information includes at least one of the purchase amount, quantity, purchase time, and purchase store.
[0010] According to the time-characteristic data set, a plurality of curvelet coefficients are obtained by curvelet transformation in combination with the curvelet characteristic frequency, specifically: According to the time-feature data set, curvelet coefficients at 6 scales are obtained through curvelet transformation; Combined with the characteristic frequency of the curvelet, the curvelet coefficients at six scales are divided into low-frequency curvelet coefficients, medium-frequency curvelet coefficients, high-frequency curvelet coefficients and ultra-high-frequency curvelet coefficients; The intermediate frequency curvelet coefficient and the high frequency curvelet coefficient both include curvelet coefficients at two scales.
[0011] Based on the low-frequency curvelet coefficients, intermediate-frequency curvelet coefficients, and high-frequency curvelet coefficients, multiple characteristic information is obtained through eigen-decomposition, specifically: Based on the low-frequency curvelet coefficients, trend characteristic information is obtained through inverse curvelet transform; Based on the intermediate-frequency curvelet coefficients, through classifying multiple directions in the intermediate-frequency curvelet coefficients and performing inverse curvelet transform, intermediate-frequency characteristic information in 4 directions is obtained; Based on the high-frequency curvelet coefficients, through classifying multiple directions in the high-frequency curvelet coefficients and performing inverse curvelet transform, high-frequency characteristic information in 4 directions is obtained.
[0012] Based on the multiple pieces of characteristic information, multiple convolution operation results are obtained through a convolutional neural network, specifically: Based on the trend characteristic information, a convolution operation result is obtained through a sharpening convolution kernel; Based on the intermediate-frequency and high-frequency characteristic information in the horizontal direction, by using a vertical edge detection convolution kernel to extract the longitudinal gradient change, convolution operation results are respectively obtained; Based on the intermediate-frequency and high-frequency characteristic information in the vertical direction, by using a horizontal edge detection convolution kernel to extract the horizontal gradient change, convolution operation results are respectively obtained; Based on the intermediate-frequency characteristic information in the diagonal direction, by using an identity convolution kernel, a convolution operation result consistent with the diagonal intermediate-frequency characteristic information is obtained; Based on the high-frequency characteristic information in the diagonal direction, by using a convolution kernel for image smoothing, a convolution operation result is obtained.
[0013] A pooling operation result is obtained through a pooling operation, specifically: Based on the multiple convolution operation results, through average pooling, a pooling operation result representing the average value within the pooling window is obtained to reduce the data dimension.
[0014] Alcohol consumption prediction is obtained through linear rectification activation, specifically: Based on the pooling operation result, combined with a weight matrix and a bias vector, an alcohol consumption prediction result is obtained; Wherein the weight matrix is a two-dimensional matrix composed of the number of output neurons and the number of input neurons, and the bias vector is a one-dimensional array of the number of output neurons, and the weight matrix and the bias vector are obtained through training.
[0015] The weight matrix and the bias vector are obtained through training, specifically: Determine training data and test data, and both the training data and the test data include input data such as the user's basic profile, daily behavior trajectory, purchase information results of alcohol consumption users, and hierarchical data obtained based on the purchase data of the current year; Performing training based on the input data and the classification data in the training data to obtain multiple sets of the weight matrices and the bias vectors; Based on the input data in the test data, multiple groups of the weight matrices and the bias vectors are combined to obtain multiple prediction results, and the weight matrix and the bias vector are determined based on the correlation between the multiple prediction results and the hierarchical data in the test data.
[0016] The present invention also provides an electronic device, comprising: Memory, for storing computer instructions; A processor is used to implement the wine consumption prediction method combining curvelet transform and convolutional neural network when executing the computer instructions.
[0017] Due to the adoption of the above technical solution, the beneficial effects achieved by the present invention are as follows: 1. In the present invention, a time-feature dataset is generated based on the user profiles, daily behavior trajectories, and purchase information of alcohol consumers. Based on this time-feature dataset, a curvelet transform is performed, combined with the curvelet characteristic frequencies, to generate multiple curvelet coefficients. These curvelet coefficients include low-frequency curvelet coefficients, mid-frequency curvelet coefficients, high-frequency curvelet coefficients, and ultra-high-frequency curvelet coefficients. By constructing an ordered model of consumer basic data, behavioral data, and consumption data, and treating this data as characteristic signals that vary over time, the present invention can more comprehensively capture consumer behavioral patterns.
[0018] Specifically, user profiles, daily behavior trajectories, and purchase results are integrated into a time-feature dataset. This approach encompasses not only static consumer attributes but also dynamic behaviors, providing a more comprehensive and in-depth understanding of consumer behavior.
[0019] Furthermore, using the curvelet transform to decompose the raw data features reveals overall trends expressed by low-frequency signals and detailed changes expressed by high-frequency signals. The curvelet transform can process data at multiple scales and angles, enhancing the data representation capabilities at each scale. For example, low-frequency components reflect the overall trend of consumer behavior, while mid-frequency and high-frequency components depict fine-grained features in different directions. This multi-level feature extraction approach enables the model to more accurately capture the implicit information in the data, including potential changes in consumer trends and behavioral patterns.
[0020] Therefore, by combining comprehensive data collection and curvelet transform technology, the present invention significantly improves the accuracy of liquor consumption prediction. This method can not only identify the basic behavior patterns of consumers, but also detect subtle behavior changes, providing more accurate decision-making support for enterprises. The present invention enhances the prediction accuracy and ensures that the model can adapt to the complex and changeable market environment.
[0021] 2. In the present invention, according to the low-frequency curvelet coefficients, medium-frequency curvelet coefficients, and high-frequency curvelet coefficients, a plurality of feature information is obtained through eigenvalue decomposition; according to the plurality of feature information, a plurality of convolution operation results are obtained through a convolutional neural network, a pooling operation result is obtained through a pooling operation, and liquor consumption prediction is obtained through linear rectification activation. The method of the present invention for enhancing feature extraction through multi-scale eigenvalue decomposition and convolutional neural network significantly enhances the stability and robustness of the liquor consumption prediction model.
[0022] First, the data is decomposed into low-frequency, medium-frequency, high-frequency, and ultra-high-frequency components by using curvelet transform, and eigenvalue decomposition is performed on these components. This decomposition method ensures that the model performs more stably in different frequency ranges: the low-frequency component reflects the overall trend, while the medium-frequency and high-frequency components respectively depict the fine features in different directions. For example, the low-frequency component captures the overall trend of consumer behavior, while the medium-frequency and high-frequency components can identify the subtle changes and trends in specific directions in the behavior pattern. This multi-level feature extraction not only helps to improve the generalization ability of the model, but also ensures its more robust performance when facing complex data.
[0023] Furthermore, the present invention uses a convolutional neural network to further extract local features, thereby effectively mining the hidden information in the feature data. This method can accurately capture the feature changes at different scales, improving the accuracy and consistency of the model in processing detailed information.
[0024] By combining the multi-scale eigenvalue decomposition of curvelet transform and the powerful feature extraction ability of convolutional neural network, the present invention significantly improves the stability and robustness of the model, which can not only adapt to different data features and application scenarios, but also effectively reduce noise interference and improve the reliability of the model in practical applications.
[0025] 3. In the present invention, a pooling operation result is obtained through a pooling operation, and liquor consumption prediction is obtained through linear rectification activation. The present invention significantly improves the intelligent level of the liquor consumption prediction model through the adaptive learning ability and efficient data processing method. First, in terms of the adaptive learning ability, the convolutional neural network can automatically adjust parameters during the training process to minimize the loss function, realizing adaptive learning. This mechanism not only improves the fitting ability of the model to complex data patterns, but also enhances its ability to respond to the changing market demands in real time.
[0026] Furthermore, the present invention utilizes pooling operations to reduce data dimensionality, thereby improving computational efficiency and preventing overfitting. By reducing the spatial size of feature maps, pooling reduces the computational complexity of subsequent layers while preserving key feature information. Through these methods, the model can maintain high accuracy while reducing computational resource consumption and improving overall operational efficiency.
[0027] In summary, this invention significantly improves the intelligence of the alcohol consumption prediction model through adaptive learning capabilities and efficient pooling operations. Adaptive learning enables the model to dynamically adjust to changing market demands, while pooling effectively reduces computational complexity and prevents overfitting, ensuring the model's stability and reliability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 The figure is a flow chart of the method for predicting alcohol consumption by combining curvelet transform and convolutional neural network according to one embodiment of the present invention. DETAILED DESCRIPTION
[0029] In order to more clearly illustrate the overall concept of the present invention, a detailed description is given below in an exemplary manner in conjunction with the accompanying drawings.
[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0031] like Figure 1 As shown in FIG, a method for predicting alcohol consumption by combining curvelet transform and convolutional neural network includes: S100: Obtain a time-feature dataset based on the user basic profile, daily behavior trajectory, and purchase information results of the alcohol consumption user.
[0032] The primary goal of this step is to integrate multi-dimensional consumer data (basic profiles, behavioral trajectories, and purchase outcomes) to construct a dataset that reflects temporal dynamics and feature correlations. This provides the foundation for subsequent feature decomposition and prediction using curvelet transforms and convolutional neural networks. The core of this phase is to transform discrete static data and dynamic behavioral data into structured data that can be analyzed by the model, thereby fully capturing the underlying patterns of consumer behavior.
[0033] It is understandable that the time - feature dataset is a structured dataset used to organize consumer behavior data in terms of time dimension and feature dimension, which can transform discrete raw data into spatio - temporal joint features that can be parsed by the model, facilitating subsequent analysis.
[0034] In this step, first, data collection and integration are carried out. User basic profiles, daily behavior trajectories, and purchase information results are collected. The above - mentioned three types of data are aligned by user ID to form a user - centered associated dataset. For example, "age 25 (basic profile)" of user A, "browse the red wine page at 8 pm every Wednesday (behavior trajectory)", and "purchase 2 bottles of wine per month (purchase result)" form an associated record.
[0035] Meanwhile, the data is sorted by timestamp to construct continuous records in the time dimension. For example, "browse the spirits page for 10 minutes on March 1st" and "purchase 1 bottle of whisky on March 15th" of user B are marked as behavior nodes on the same timeline.
[0036] It should be noted that data dimension mapping and standardization are required. Feature encoding is performed to convert unstructured data (such as text descriptions) into numerical features. For example, gender is encoded as binary (0 / 1), and region is encoded as region ID. At the same time, normalization processing is carried out to standardize numerical data (such as purchase amount, browsing duration) (such as Z - score normalization) to eliminate the dimension difference.
[0037] During this process, feature weights are dynamically assigned, and different weights are given according to the importance of behaviors. For example, the weight of purchase amount is higher than that of browsing duration to strengthen the impact of key behaviors on prediction. In addition, missing data is filled by interpolation or mean to ensure data integrity.
[0038] Finally, a time - feature dataset is constructed. The processed data is organized into a three - dimensional tensor (user × time × feature), where: user, each user is an independent sample; time, divided by a fixed time window (such as week / day) to form time steps; feature, including the encoded values of basic profiles, behavior trajectories, and purchase results.
[0039] A sliding time window (such as the data in the recent 30 days) is used to dynamically update features to capture the time - variability of behavior patterns. For example, "the browsing frequency of user C has increased in the recent 7 days" may indicate an enhanced purchase intention.
[0040] In this step, the data is aligned by the user ID to avoid prediction biases caused by isolated analysis of a single dimension (such as only focusing on the purchase amount while ignoring behavioral changes). The prediction weight of key behaviors (such as purchase records) is strengthened, interference from redundant information is reduced, and the feature decomposition efficiency of the subsequent curvelet transform is improved. The sliding window balances the influence of historical behaviors and recent trends, enhancing the model's sensitivity to short-term behavioral changes (such as consumption fluctuations during holidays).
[0041] This step lays the foundation for the multi-scale feature decomposition of the subsequent curvelet transform and the local feature extraction of the convolutional neural network through structured data integration, which is a key prerequisite for improving prediction accuracy and model stability.
[0042] S200: According to the time-feature data set, through curvelet transform and in combination with curvelet characteristic frequencies, a plurality of curvelet coefficients are obtained. The plurality of curvelet coefficients include low-frequency curvelet coefficients, medium-frequency curvelet coefficients, high-frequency curvelet coefficients, and ultra-high-frequency curvelet coefficients.
[0043] In this step, the time-feature data set is subjected to multi-scale and multi-directional feature decomposition through curvelet transform, and the original data is converted into low-frequency, medium-frequency, high-frequency, and ultra-high-frequency curvelet coefficients. This process can separate the global trend (low frequency) and local details (high frequency) in the data, providing a clearer and more structured input for the subsequent feature extraction of the convolutional neural network, thereby improving the accuracy and robustness of the prediction model.
[0044] It can be understood that curvelet transform is a signal processing technology based on multi-scale and multi-directional analysis, which is specifically for the efficient decomposition of curvilinear singularities in two-dimensional data (such as images or time-feature data). It can cover trends in different time ranges (such as monthly, weekly, daily) through the scale parameter, and capture the directional features in the data (such as the curve changes in the user browsing trajectory) through the direction parameter.
[0045] In the present invention, curvelet transform solves the problem that traditional methods (such as wavelet transform) cannot effectively analyze curvilinear consumption behaviors (such as holiday consumption patterns).
[0046] Curvelet coefficients refer to the feature vectors obtained after curvelet transform. Each coefficient corresponds to the data energy distribution at a specific scale (frequency) and direction. Among them, the low-frequency coefficients represent the global trend (such as the user's average monthly consumption level). The medium-high frequency coefficients capture local details (such as a sudden increase in daily consumption).
[0047] By classifying the curvelet coefficients, the model can specifically extract features at different levels and avoid information aliasing.
[0048] In this step, the parameter configuration of the curvelet transform is preset, and the scale parameter and direction parameter of the curvelet transform are determined to control the fineness of decomposition. The low-frequency components correspond to large scales, and the high-frequency components correspond to small scales, ensuring coverage of all frequency bands from the overall trend to the subtle changes. Additionally, high-frequency components (such as intermediate frequency and ultra-high frequency) require more direction parameters to capture complex behavior patterns (such as the curve changes in the user browsing trajectory).
[0049] After the preset of the curvelet transform algorithm, the time-feature dataset (such as the three-dimensional tensor of user D) is input into the curvelet transform algorithm. The data is decomposed into subbands of different scales through a recursive filter bank. Directional decomposition is performed on each subband to generate curvelet coefficients in different directions.
[0050] According to the curvelet characteristic frequencies, the coefficients are divided into four categories: low-frequency curvelet coefficients, covering the global trend (such as the average monthly consumption amount of users); intermediate-frequency curvelet coefficients, capturing medium details (such as weekly consumption fluctuations); high-frequency curvelet coefficients, depicting short-term behavior changes (such as a sudden increase in the single-day browsing frequency); ultra-high-frequency curvelet coefficients, filtering noise (such as data acquisition errors).
[0051] By dynamically adjusting the scale and direction parameters, the model can adapt to the data characteristics of different users (such as users with frequent high-frequency behaviors requiring more direction parameters). The threshold setting of the ultra-high-frequency coefficients ensures that the model focuses on meaningful features and avoids noise interfering with the prediction results.
[0052] Through the above steps, the curvelet transform converts the original data into structured multi-scale coefficients, providing a clear and low-noise input for the subsequent feature extraction of the convolutional neural network, significantly improving the accuracy and stability of the prediction model. This process connects data preprocessing with the deep learning model to achieve the final consumption prediction.
[0053] S300: According to the low-frequency curvelet coefficients, intermediate-frequency curvelet coefficients, and high-frequency curvelet coefficients, multiple feature information is obtained through feature decomposition.
[0054] The main purpose of this step is to extract fine-grained feature information in different frequency bands by further decomposing the low-frequency, intermediate-frequency, and high-frequency curvelet coefficients, convert the multi-scale features of the original data into structured feature vectors that can be effectively utilized by the convolutional neural network, and enhance the model's expression ability for complex behavior patterns. This process can isolate features strongly related to consumption prediction (such as long-term trends, short-term fluctuations, behavior mutations, etc.) and eliminate redundant information, providing an efficient input for subsequent prediction.
[0055] It should be noted that the low-frequency, intermediate-frequency, and high-frequency curvelet coefficients have different physical meanings. For example, the low-frequency curvelet coefficients contain overall trend information, and the corresponding decomposition algorithm needs to be selected according to the physical meaning of the curvelet coefficients.
[0056] For example, principal component analysis (PCA) is used to extract core trend features for low-frequency coefficients. Independent component analysis (ICA) is adopted for intermediate-frequency coefficients to separate directional behavior patterns (such as the curve direction of the user's browsing trajectory). For high-frequency coefficients, wavelet packet decomposition is used to further subdivide high-frequency components and capture sudden behavior changes.
[0057] The optimal algorithm is selected for different frequency bands (e.g., PCA is suitable for dimensionality reduction trends, and wavelet packets are suitable for high-frequency details). At the same time, features with high contribution to prediction are screened out through analysis of variance (ANOVA), and redundant components are removed.
[0058] According to the corresponding algorithms, feature information extraction and encoding are carried out. Among them, for low-frequency feature extraction, PCA dimensionality reduction is performed on the low-frequency coefficient matrix to generate a feature vector reflecting long-term consumption ability (such as the "monthly average consumption stability index"). For intermediate-frequency feature extraction, ICA decomposition is performed on the intermediate-frequency coefficients to separate the directional features of user behavior (such as "horizontal browsing preference" or "vertical click hotspot"). For high-frequency feature extraction, wavelet packet decomposition is performed on the high-frequency coefficients to generate features reflecting short-term behavior mutations (such as "sudden increase in the number of daily browsing times").
[0059] Semantic labels are assigned to each decomposed feature (such as "high frequency - browsing duration volatility") for subsequent model interpretation. Different weights are assigned according to the feature contribution degree (e.g., the weight of high-frequency features is higher than that of low-frequency ones to highlight the impact of short-term behavior changes).
[0060] Finally, feature fusion and storage are carried out. The low-frequency, intermediate-frequency, and high-frequency feature vectors are concatenated into a multi-scale feature matrix. The feature dimensions are unified through standardization (such as Z-score) to form the final input feature information.
[0061] During the feature fusion and storage process, weighted summation or concatenation methods are adopted to ensure the complementarity of features in different frequency bands. At the same time, highly correlated features are removed through correlation analysis to reduce computational redundancy.
[0062] Through this step, the original curvelet coefficients are transformed into structured and low-redundancy multi-scale feature information, providing accurate input for the local feature extraction of the subsequent convolutional neural network and significantly improving the model's ability to capture complex consumption behaviors.
[0063] S400: According to the multiple pieces of feature information, through a convolutional neural network, multiple convolution operation results are obtained, a pooling operation result is obtained through pooling operation, and liquor consumption prediction is obtained through linear rectification activation.
[0064] The main purpose of this step is to perform local feature extraction, dimensionality reduction, and non - linear modeling on multi - scale feature information through a Convolutional Neural Network (CNN), and finally output the prediction result of alcohol consumption. The core of this step is to utilize the local perception ability of the CNN to capture the spatial / temporal associations between features, reduce overfitting through pooling, and enhance the model's fitting ability for complex patterns through activation functions, so as to achieve high - precision and stable consumption prediction.
[0065] It can be understood that a Convolutional Neural Network (CNN) is a deep - learning model that processes data with a grid structure (such as images, time series) through convolutional layers, pooling layers, and fully - connected layers. Its core lies in that the convolutional kernel only perceives a local area, reducing the number of parameters and enhancing the ability to capture spatial / temporal associations. The same convolutional kernel shares parameters at different positions, improving generalization.
[0066] Convolution operation (Convolution Operation) is to perform a dot - product operation between the sliding convolutional kernel and the input data to extract local features. In the present invention, the convolutional layer performs pattern matching on multi - scale features.
[0067] Pooling operation (Pooling Operation) reduces the spatial dimension of the feature map through downsampling while retaining key information. Max - pooling is used to retain local maxima, enhancing the robustness to significant features. Average - pooling is used to smooth noise and improve the stability of low - frequency features.
[0068] The Rectified Linear Unit (ReLU) is a non - linear activation function that avoids the vanishing gradient problem, accelerates model convergence, and reduces computational complexity through sparse activation.
[0069] In this step, multi - scale feature information (such as low - frequency, medium - frequency, and high - frequency feature vectors) is concatenated into a three - dimensional tensor (feature × time × user) as the input of the CNN. Multiple groups of convolutional kernels are designed, and optimal parameters are selected for different feature types. The size and stride of the convolutional kernel are dynamically adjusted according to the feature frequency. For example, a stride of 1 is used for high - frequency features to retain details. Additionally, low - frequency, medium - frequency, and high - frequency features are input into different convolutional branches to avoid information aliasing.
[0070] In addition, convolution operation and local feature extraction are performed. Convolution operations are carried out on the input tensor to generate multiple feature maps. For example, for the low - frequency convolutional branch, a "long - term trend feature map" (such as consumption stability) is output. For the high - frequency convolutional branch, a "short - term fluctuation feature map" (such as a sudden increase in browsing frequency) is output.
[0071] In this step, through convolutional kernel initialization, the He initialization method is adopted to avoid gradient vanishing / explosion. Different frequency band features are integrated through cross-branch connections (such as residual connections) for feature fusion to enhance the model's expressive ability.
[0072] Again, perform pooling operations and dimensionality reduction. Apply a pooling layer to each feature map to reduce the spatial dimension (such as the time step or feature dimension). For example, max pooling retains the local maximum value to highlight significant features (such as "short-term interest surges"). Average pooling smooths the noise and enhances the stability of low-frequency features (such as "monthly average consumption").
[0073] Select the pooling type according to the feature importance (such as using max pooling for high-frequency features and average pooling for low-frequency features). The pooling window is set to 2×2 or 3×3 to balance the dimensionality reduction efficiency and information retention.
[0074] Finally, perform the fully connected layer and activation function. Flatten the pooled features and input them into the fully connected layer for final prediction. Use the rectified linear unit (ReLU) activation function to output the result.
[0075] It should be noted that a Dropout layer (such as a retention rate of 0.5) is added before the fully connected layer to prevent overfitting. If it is necessary to predict both the purchase probability and amount simultaneously, the output nodes can be expanded and a multi-objective loss function can be designed.
[0076] This step significantly improves the comprehensive performance of liquor consumption prediction through the deep learning architecture of the convolutional neural network: First, the convolutional layer accurately captures the local correlations of multi-scale features (such as the user's long-term consumption ability and short-term behavior mutations). Combined with the dimensionality reduction processing of the pooling operation, the model can identify complex behavior patterns (such as the superposition effect of "short-term interest surges × long-term consumption ability"); Second, the pooling operation effectively suppresses the risk of overfitting, ensuring that the model can still maintain stable performance in the scenarios of new users or sparse data; In addition, the weight sharing mechanism of the convolutional layer and the dimensionality compression ability of the pooling layer reduce the number of model parameters, significantly improving the computational efficiency and inference speed; Finally, through the ReLU activation function, the model can fit the complex relationships in consumption behavior (such as the multiplier effect of "promotion day + interest surge"), further enhancing the robustness and generalization ability of the prediction.
[0077] Generally speaking, in this step, through the deep learning ability of the CNN, multi-scale features are transformed into accurate prediction results, solving the problems of inaccurate prediction and unstable model caused by insufficient feature correlation in traditional methods, and providing an efficient and reliable solution for liquor consumption prediction.
[0078] As a preferred implementation manner of the present invention, according to the user basic profile, daily behavior trajectory, and purchase information results of liquor consumption users, a time-feature data set is obtained, specifically: According to the user base profile, daily behavior trajectory, and purchase information results of liquor consumption users, perform formatting transformation to obtain time series values one by one corresponding to them. Obtain a time-feature dataset according to the time series values.
[0079] The main purpose of this embodiment is to integrate the user base profile, daily behavior trajectory, and purchase information results into a structured time-feature dataset, providing high-quality input for subsequent wavelet transform and convolutional neural network analysis. This process transforms the original data into a unified time series form through formatting transformation, ensuring the alignment and parsability of multi-dimensional data in the time dimension.
[0080] First, formatting transformation and construction of time series values. Extract the original data from the user base profile, daily behavior trajectory, and purchase information results. Convert unstructured data (such as text, dates) into numerical forms: In the user base profile, age (numerical value), gender (binary encoding, such as 0 / 1), date of birth (calculated from age), registration time (timestamp).
[0081] In the daily behavior trajectory, login time (timestamp), browsing commodity time (duration), number of order cancellations (count), number of favorited commodities (count).
[0082] In the purchase result information, purchase amount (numerical value), quantity (count), purchase time (timestamp), purchase store (categorical encoding, such as 1 = online, 2 = offline).
[0083] Align all data according to a unified time granularity (such as daily, weekly) to generate time series values.
[0084] Among them, t represents the time point, N represents the selectable consumption features, y represents the consumption feature N at time t in numerical expression.
[0085] In the formatting transformation and construction of time series values, static data (such as age) is repeatedly filled to ensure that all features are consistent in the time dimension. For example, the age of user C, which is 30 years old, is 30 at all time points. At the same time, missing value processing is performed, and missing data is filled by interpolation (such as linear interpolation) or mean value to avoid data breaks.
[0086] Secondly, construction of the time-feature dataset. Organize the formatted data into a three-dimensional tensor (user × time × feature) according to the time dimension and feature dimension. The specific form is:
[0087] in: Represents a collection of profile characteristics (e.g., age, gender); Represents a collection of behavioral characteristics (such as browsing time, number of collections); Represents a set of purchase characteristics (such as purchase amount and quantity).
[0088] Defining ranking based on feature importance S , Indicates that the archive data is sorted in ascending order by timestamp. Indicates that the behavior data is sorted in descending order by the frequency of the behavior; Indicates that the purchase data is sorted in descending order by purchase amount.
[0089] This implementation aligns data by user ID, ensuring that all features for the same user are consistent across time. Using sliding window technology, a fixed window (e.g., the last 30 days) is used to dynamically update the dataset, capturing the time-varying nature of behavioral patterns.
[0090] This implementation ensures data consistency and integrity through formatting and time series mapping. For example, a user's "age 30" remains consistent across all time points, avoiding bias in isolated analysis and improving data quality. By aligning user IDs, the model can identify the correlation between a user's "high-end red wine preference" (behavioral trajectory) and "average monthly spending of 800 yuan" (purchase data), enhancing multi-dimensional correlation. Sliding window technology can detect user C's "recent surge in browsing frequency," capturing dynamic behavior and providing time-sensitive features for prediction.
[0091] As a preferred embodiment of this embodiment, the basic user profile, daily behavior trajectory, and purchase information results of the alcohol consumption user are specifically as follows: The basic user profile includes at least one of the consumer's age, gender, date of birth, and registration time; Daily behavior traces include at least one of the following: time logged into the platform, time spent browsing products, time initiating payment, number of order cancellations, number of product collections, and number of activity participations; The purchase result information includes at least one of the purchase amount, quantity, purchase time, and purchase store.
[0092] The main purpose of this embodiment is to clarify the specific content of user basic files, daily behavior trajectories and purchase result information, ensure that the data covers the full-dimensional characteristics of consumer behavior, and provide quantifiable input for subsequent analysis.
[0093] The user's basic profile includes at least any one of the following fields: age, gender, date of birth, registration time. It should be noted that for static feature encoding, gender is encoded as binary (0 / 1), date of birth is converted to age, and registration time is converted to registration duration (such as "registered for 1 year").
[0094] It should be noted that if the data permits, other fields (such as occupation, region) can be added, and the present invention does not limit this.
[0095] The daily behavior trajectory includes at least any one of the following fields: time of logging in to the platform, time of browsing products, time of initiating payment, number of order cancellations, number of products favorited, number of activity participations.
[0096] In the daily behavior trajectory data, for behavior quantification, the "time of logging in to the platform" is recorded as a timestamp, and the "time of browsing products" is converted to a duration (such as "staying for an average of 15 minutes each time"). At the same time, behavior clustering is performed, combining the "time of initiating payment" and the "number of order cancellations" to identify users' payment hesitation behavior.
[0097] The purchase result information includes at least any one of the following fields: purchase amount, quantity, purchase time, purchase store.
[0098] For purchase pattern analysis, combining the "purchase amount" and the "purchase time" to identify the consumption peak on promotional days; encoding the "purchase store" as a categorical variable (such as online / offline). At the same time, for outlier handling, a threshold is set for the "purchase amount" (such as excluding outliers > 100,000 yuan) to reduce noise interference.
[0099] In this embodiment, to ensure feature comprehensiveness, combining the "age" in the user's basic profile with the "amount" in the purchase result can identify the correlation between age and consumption ability (such as a 30-year-old user with an average monthly consumption of 800 yuan). To identify behavior patterns, combining the "number of favorites" in the daily behavior trajectory with the "quantity" in the purchase result can judge the user's interest and purchase conversion rate (such as purchasing 2 bottles after favoriting 10 products). For scenario analysis, the "online / offline" classification of the purchase store can analyze channel preferences. For example, users prefer to purchase high-end red wine online.
[0100] Specifically, the user's data is as follows: Basic profile, age 35 years old (numerical value), gender female (encoded as 1), registration time January 2020 (registration duration 4 years).
[0101] Behavior trajectory, logging in to the platform 3 times a week, staying for an average of 10 minutes. The time spent browsing the red wine page accounts for 60%, and the number of products favorited per month is 5 times.
[0102] Purchase results: I bought red wine 3 times in the past six months, with an average single purchase amount of 600 yuan, and all purchases were made online.
[0103] Based on the above data, a time-feature dataset was constructed. The profile characteristics are: age 35, gender 1, registration duration 4 years; behavioral characteristics: login 3 times a week, number of collections 5 times / month; purchase characteristics: purchase amount 600 yuan / time, online purchase.
[0104] By associating the "number of collections" in the behavioral trajectory with the "amount" of the purchase result, we can predict user D's potential purchase intention for high-end red wine.
[0105] Overall, this implementation improves data integrity and consistency. Formatting conversion and time series mapping ensure alignment of all features across the time dimension, avoiding biases that can arise from isolated analysis. By integrating multi-dimensional features, user profiles, behavioral trajectories, and purchase results, we can capture the correlation between long-term trends (such as age and spending power) and short-term behaviors (such as a surge in favorites). This provides scenario-based forecasting, and the classification coding of purchasing stores supports channel preference analysis, providing accurate consumption forecasts.
[0106] This implementation constructs a structured, high-dimensional time-feature dataset, laying a solid foundation for feature decomposition and prediction using curvelet transform and convolutional neural networks, significantly improving the accuracy and generalization ability of alcohol consumption prediction.
[0107] As a preferred embodiment of the present invention, based on the time-feature data set, a plurality of curvelet coefficients are obtained by curvelet transformation in combination with the curvelet characteristic frequency, specifically: According to the time-feature data set, curvelet coefficients at 6 scales are obtained through curvelet transformation; Combined with the characteristic frequency of the curvelet, the curvelet coefficients at six scales are divided into low-frequency curvelet coefficients, medium-frequency curvelet coefficients, high-frequency curvelet coefficients and ultra-high-frequency curvelet coefficients; The intermediate frequency curvelet coefficient and the high frequency curvelet coefficient both include curvelet coefficients at two scales.
[0108] The main purpose of this implementation is to decompose a time-feature dataset into six-scale curvelet coefficients using the curvelet transform. These coefficients are then classified into four frequency categories: low-frequency, medium-frequency, high-frequency, and ultra-high-frequency. The medium-frequency and high-frequency categories each contain two scales. This process aims to separate global trends (low-frequency), medium-level details (medium-frequency), short-term fluctuations (high-frequency), and noise (ultra-high-frequency) in the data, providing clear multi-scale features for subsequent feature decomposition.
[0109] The scale corresponds to the coarseness or fineness of the frequency. A low scale (such as Scale 1) corresponds to low frequencies (global trends), and a high scale (such as Scale 6) corresponds to high frequencies (details or noise).
[0110] The low-frequency curve coefficient (C{1}) corresponds to global trends, such as users' long-term spending power and registration duration.
[0111] The medium-frequency curvature coefficients (C{2}, C{3}) correspond to medium-level details, such as weekly consumption fluctuations and periodicity in activity participation. The high-frequency curvature coefficients (C{4}, C{5}) correspond to short-term mutations, such as a single-day surge in browsing and shifts in interest. The ultra-high-frequency curvature coefficient (C{6}) corresponds to noise or outliers, such as data collection errors or misoperation.
[0112] Input the time-feature dataset (such as the three-dimensional tensor of user X) into the curvelet transform algorithm and perform the curvelet transform to obtain the curvelet coefficients at different scales and angles.
[0113] in, Indicated on scale ,direction ,Location The curvelet coefficient on Indicates any sort order The feature data set, Indicates consumption data in feature dimension and time The distribution of represents the curvelet function.
[0114] Perform multi-scale decomposition and decompose the data into 6 scales (Scale 1 to Scale 6) through a recursive filter bank: Scale 1 represents low-frequency components (global trends); Scale 2-3 represents medium-frequency components (medium details); Scale 4-5 represents high-frequency components (short-term mutations); Scale 6 represents ultra-high-frequency components (noise).
[0115] Directional parameter setting: the mid-frequency scale (Scale 2-3) uses 32 directional parameters, and the high-frequency scale (Scale 4-5) uses 64 directional parameters, covering the geometric directions of two-dimensional space (such as 0°, 45°, 90°, etc.).
[0116] The numerical distribution of the curvelet coefficient set is shown in the following table:
[0117] For low frequencies (Scale 1), only isotropic features are retained. For medium and high frequencies (Scale 2 - 5), directional analysis is required. A threshold is set for the ultra-high frequency coefficients (Scale 6) (e.g., retain coefficients with absolute value > 0.1) to filter out noise.
[0118] Perform frequency classification and coefficient categorization. For low-frequency coefficients, the coefficients of Scale 1 are categorized as low-frequency curvelet coefficients (C{1}). For medium-frequency coefficients, the coefficients of Scale 2 and 3 are categorized as medium-frequency curvelet coefficients (C{2}, C{3}). For high-frequency coefficients, the coefficients of Scale 4 and 5 are categorized as high-frequency curvelet coefficients (C{4}, C{5}). For ultra-high frequency coefficients, the coefficients of Scale 6 are categorized as ultra-high frequency curvelet coefficients (C{6}).
[0119] For medium frequencies (C{2}, C{3}) and high frequencies (C{4}, C{5}), all directional information needs to be retained to support subsequent directional feature extraction. The directional weights of the high-frequency coefficients (C{4}, C{5}) are adjusted according to behavioral importance (e.g., higher weight for "diagonal browsing").
[0120] Specifically, the user's original data: low frequency, monthly average consumption of 800 yuan (long-term trend); medium frequency, browsing the wine page for 20 minutes every Wednesday (weekly pattern); high frequency, a 300% surge in browsing high-end products on the promotion day (short-term mutation); noise, accidental cancellation of an order (Scale 6).
[0121] The results after decomposition are: C{1} (low frequency), extracting the trend of "monthly average consumption of 800 yuan" for predicting long-term repurchase potential; C{2} (medium frequency), the direction parameter identifies the "Wednesday browsing preference" (horizontal direction); C{4} (high frequency), the direction parameter captures the "diagonal browsing of high-end products" (diagonal direction); C{6} (ultra-high frequency), filtering out the noise of "accidental cancellation of an order" to avoid misjudging the purchase intention.
[0122] In this embodiment, multi-scale feature separation, the separation of low frequency (C{1}) and high frequency (C{4} - C{5}) enables the model to capture both "long-term stable users" and "short-term interest surges" behaviors simultaneously, and the prediction accuracy is increased to 85%. Enhancing directional details, the direction parameters of medium frequency (C{2} - C{3}) and high frequency (C{4} - C{5}) (such as 32 directions) support the geometric feature analysis of behavioral patterns (e.g., "diagonal browsing" corresponding to the interest in high-end products). Noise suppression and improved robustness, the threshold filtering of ultra-high frequency (C{6}) reduces the interference of outliers and ensures the stability of the model in data-sparse scenarios (such as new users).
[0123] This implementation uses multi-scale, multi-directional decomposition using the curvelet transform to convert the time-feature dataset into four types of coefficients: low-frequency, medium-frequency, high-frequency, and ultra-high-frequency. This provides structured, high-signal-to-noise ratio input features for the subsequent convolutional neural network. This not only improves the model's ability to capture complex patterns in consumer behavior but also enhances the robustness of predictions through noise filtering, laying a key foundation for accurate consumer forecasting.
[0124] As a preferred embodiment of this implementation, multiple feature information is obtained by feature decomposition based on the low-frequency curvelet coefficient, the intermediate-frequency curvelet coefficient, and the high-frequency curvelet coefficient, specifically: According to the low-frequency curvelet coefficient, the trend characteristic information is obtained through the inverse curvelet transform; According to the intermediate frequency curvelet coefficient, the intermediate frequency curvelet coefficient is classified into multiple directions, and the intermediate frequency feature information in four directions is obtained through the inverse curvelet transform; According to the high-frequency curvelet coefficients, multiple directions in the high-frequency curvelet coefficients are classified, and the high-frequency feature information in four directions is obtained through inverse curvelet transformation.
[0125] This implementation aims to separate the low-frequency, mid-frequency, and high-frequency features of consumer data into interpretable directional information through coefficient decomposition and inverse transformation of the curvelet transform. Trend features are extracted, and the overall trend of consumer data (such as long-term spending power) is obtained through inverse transformation of the low-frequency coefficients. Directional features are decoupled, and through directional classification and inverse transformation of mid-frequency and high-frequency data, fine-grained features in different directions (such as horizontal browsing and diagonal clicks) are separated, improving the model's ability to capture complex behavioral patterns. Sparse representation optimization is performed, leveraging the sparsity of the curvelet coefficients to reduce redundant information and improve computational efficiency.
[0126] For the low-frequency curvelet coefficients, coefficient filtering is performed to retain only the low-frequency curvelet coefficients (C{1}), and the coefficients of other scales (C{2}-C{6}) are set to zero.
[0127] Restore global trend characteristics.
[0128] in, Represents the trend information of consumption feature data in the time-feature domain, represents the curvelet coefficient of the low-frequency component, represents the inverse curvelet transform.
[0129] This embodiment utilizes the sparsity of low-frequency coefficients to retain only significant coefficients (e.g., absolute value > 0.1) to improve computational efficiency. The inverse transform result is low-pass filtered and smoothed to eliminate short-term fluctuation interference.
[0130] For the intermediate frequency direction classification and inverse transformation, the intermediate frequency coefficients (C{2}, C{3}) are each divided into 32 directions, and the first 8 directions are divided into 4 groups (2 directions in each group): Group 1, directions 1-2 (C{2}{1-2}, C{3}{1-2}); Group 2, directions 3-4 (C{2}{3-4}, C{3}{3-4}); Group 3, directions 5-6 (C{2}{5-6}, C{3}{5-6}); Group 4, directions 7-8 (C{2}{7-8}, C{3}{7-8}).
[0131] Perform inverse transformation on the coefficients of each group of directions to generate intermediate frequency features in 4 directions
[0132] Among them, represents the detailed information of the consumption feature data in the first direction of the time-feature domain, represents the curvelet coefficients in the first 8 directions under the intermediate frequency component, represents the inverse curvelet transform.
[0133] Similarly, according to the above process to divide the curvelet coefficients in different directions, the characterization information of the other 3 directions corresponding to the intermediate frequency signal in the time-feature domain can be calculated, which are respectively denoted as: .
[0134] Define direction labels according to the behavior pattern (for example, Group 1 is "horizontal browsing", and Group 4 is "diagonal right collection"). Assign higher weights to the directions strongly related to consumption (such as "diagonal browsing of high-end goods").
[0135] For the high frequency direction classification and inverse transformation, perform direction grouping. The high frequency coefficients (C{4}, C{5}) are each divided into 64 directions, and the first 16 directions are divided into 4 groups (4 directions in each group): Group 1, directions 1-4 (C{4}{1-4}, C{5}{1-4}); Group 2, directions 5-8 (C{4}{5-8}, C{5}{5-8}); Group 3, directions 9-12 (C{4}{9-12}, C{5}{9-12}); Group 4, directions 13-16 (C{4}{13-16}, C{5}{13-16}).
[0136] Perform inverse transformation on the coefficients of each group of directions to generate high frequency features in 4 directions .
[0137] Perform wavelet threshold processing (such as the soft threshold method) on the high frequency coefficients to filter out small fluctuations. According to the mutation priority, retain the coefficients with high mutation energy in the direction group (such as "sharp increase in diagonal browsing on promotion days").
[0138] For ultra-high frequency noise filtering, directly filter the coefficients of C{6} (ultra-high frequency) because they correspond to noise or outliers (such as misoperations) and do not participate in the inverse transform.
[0139] Specifically, the user's original data is as follows: low frequency (C{1}), with a monthly average consumption of 1000 yuan (trend feature); medium frequency (C{2}-C{3}): Group 1 (direction 1-2), browsing the red wine page horizontally for 20 minutes / week, Group 4 (direction 7-8), diagonally rightward collecting high-end products; high frequency (C{4}-C{5}): Group 1 (direction 1-4), a 300% surge in browsing diagonally leftward on promotional days, Group 4 (direction 13-16), quickly clicking the payment button.
[0140] Execute the decomposition result to obtain the trend feature. Identify "long-term consumption stability index = 95%".
[0141] Extract the medium-frequency features. For Group 1 (X SM1 ), the duration of horizontally browsing red wine increases by 30%. For Group 4 (X SM4 ), the frequency of diagonally rightward collecting high-end products increases.
[0142] Extract the high-frequency features. For Group 1 (X SH1 ), there is a 300% surge in browsing high-end products diagonally leftward on promotional days. For Group 4 (X SH4 ), the number of times of quickly clicking the payment button doubles on promotional days.
[0143] In this embodiment, through the feature decomposition and direction classification of curvelet transform, multi-scale separation is achieved, separating the low-frequency trend, medium-frequency periodicity, and high-frequency mutation, covering the full-scene features of consumption behavior. Through directional decoupling, the direction grouping of medium frequency and high frequency (such as 4 groups of directions) ensures feature independence and improves the model's analytical ability for complex behaviors. Through sparse representation optimization, only the coefficients of key directions are retained, reducing redundant information, accelerating calculations, and improving model efficiency.
[0144] This embodiment combines the sparsity and directional analysis of curvelet transform, providing structured and high-signal-to-noise input features for the subsequent convolutional neural network, significantly improving the accuracy and interpretability of consumption prediction.
[0145] Specifically, according to the multiple feature information, through a convolutional neural network, multiple convolution operation results are obtained, specifically: According to the trend feature information, through a sharpening convolution kernel, a convolution operation result is obtained; According to the horizontal medium-frequency feature information and high-frequency feature information, through a vertical edge detection convolution kernel, the longitudinal gradient change is extracted, and convolution operation results are obtained respectively; According to the longitudinal intermediate-frequency feature information and high-frequency feature information, the horizontal gradient change is extracted through a horizontal edge detection convolution kernel, and the convolution operation results are obtained respectively; According to the oblique intermediate-frequency feature information, a convolution operation result consistent with the oblique intermediate-frequency feature information is obtained through a unit convolution kernel; According to the oblique high-frequency feature information, a convolution operation result is obtained through a convolution kernel for image smoothing.
[0146] The purpose of this embodiment is to perform differential processing on the nine-direction information of low-frequency, intermediate-frequency, and high-frequency features through a directionally designed convolution kernel, enhance effective features, suppress noise, and finally output a feature map strongly related to the consumer behavior pattern.
[0147] First, trend feature processing (low frequency) is performed. The trend feature information (X SL ), which represents the global trend of consumption data (such as monthly average consumption). The convolution kernel selects a sharpening convolution kernel ,
[0148] Through convolution operation
[0149] Among them, represents the convolution operation result.
[0150] The mutation points in the trend are highlighted through a center-weighted kernel (such as "a sharp increase in consumption during Double Eleven"), while details are retained. For example, the monthly fluctuations hidden in the "long-term consumption stability" trend of users are enhanced.
[0151] In addition, it is convenient for key point detection. For example, the "sudden drop in monthly average consumption" of user B is sharpened and more easily recognized by the model.
[0152] Process the horizontal intermediate-frequency and high-frequency features. The horizontal intermediate-frequency X SM1 (the intermediate-frequency behavior distributed horizontally, such as "weekly horizontal browsing of red wine"), and the horizontal high-frequency X SH1 (the high-frequency behavior distributed horizontally, such as "a sharp increase in horizontal browsing on promotion days").
[0153] The convolution kernel selects a vertical edge detection convolution kernel (w vt ), for example
[0154] used to detect the longitudinal gradient change of the horizontally distributed signal (such as "vertical fluctuations in browsing duration").
[0155] The convolution operation formula is expressed as
[0156]
[0157] Among them, represents the result of the convolution operation.
[0158] The gradient change in the vertical direction corresponds to the mutation of the behavior (such as the sudden increase in the browsing duration on the promotion day). The separation of the horizontal behavior and the vertical change avoids interfering with the features in other directions. For mutation recognition, for example, in the "horizontal browsing every Wednesday" trend of the user, the abnormal fluctuation during the promotion week is captured by the vertical edge detection kernel.
[0159] Process the vertical medium-frequency / high-frequency features, where the vertical medium-frequency X SM3 (the medium-frequency behavior distributed vertically, such as "the duration of scrolling the detailed page"). The vertical high-frequency X SH3 (the high-frequency behavior distributed vertically, such as "fast scrolling on the promotion day").
[0160] The convolution kernel selects the horizontal edge detection convolution kernel (w zh )), for example
[0161] is used to detect the horizontal gradient change of the vertically distributed signal (such as "the horizontal fluctuation of the scrolling duration").
[0162] The convolution operation formula is expressed as
[0163]
[0164] Among them, represents the result of the convolution operation.
[0165] The horizontal gradient analysis identifies the mutation in the vertical behavior (such as "fast scrolling the detailed page on the promotion day" of user F). The separation of the vertical behavior and the horizontal change avoids cross-direction interference. For behavior pattern recognition, for example, the sudden increase in the "vertical scrolling" behavior of the user during the high-frequency promotion period is captured by Y SH3 .
[0166] For the processing of the inclined medium-frequency features, the oblique medium-frequency X SM2 (oblique to the left), X SM4 (oblique to the right), represents the medium-frequency behavior in the inclined direction (such as "collecting goods obliquely").
[0167] The convolution kernel selects the unit convolution kernel (w id ), for example:
[0168] is used to retain the original features in the inclined direction and avoid information loss.
[0169] The convolution operation formula is expressed as
[0170]
[0171] in, Represents the result of the convolution operation.
[0172] The original data of the tilt direction is directly passed to the next layer, preserving the direction information. This ensures directional integrity. For example, a user's "diagonally collecting high-end products" behavior is fully preserved, supporting subsequent model analysis.
[0173] For tilted high frequency feature processing, tilted high frequency: X SH2 (diagonally left), X SH4 (diagonally to the right), indicating high-frequency behavior in the diagonal direction (e.g., a “surge in diagonal browsing”).
[0174] The convolution kernel selects Gaussian smoothing convolution kernel (w gs ),For example:
[0175] Used to suppress noise (such as erroneous operations) and retain valid mutations.
[0176] The convolution operation formula is expressed as
[0177]
[0178] in, Represents the result of the convolution operation.
[0179] The Gaussian smoothing kernel reduces high-frequency noise while preserving true mutation signals. This achieves noise suppression. For example, during a user's "surge in diagonal browsing," the noise generated by misoperation is smoothed while the true shift in interest is preserved.
[0180] This embodiment uses vertical and horizontal edge detection kernels to accurately extract gradient changes in horizontal or vertical behavior (e.g., "surge in browsing on promotional days"), enhancing directional features. Frequency-specific processing enhances trend details by sharpening low frequencies, suppresses noise by smoothing high frequencies, and preserves original features by tilting mid-frequency directions. Nine independent channels prevent cross-directional interference, improving the model's ability to analyze complex behavioral patterns.
[0181] This embodiment follows the feature extraction principle of convolutional neural networks and significantly improves the accuracy and robustness of the consumption prediction model through directional convolution kernel optimization.
[0182] As a preferred embodiment of the present invention, a pooling operation result is obtained through pooling operation, specifically: Based on multiple convolution operation results, through average pooling, a pooling operation result representing the average value within the pooling window is obtained to reduce the data dimension.
[0183] This embodiment aims to reduce the spatial resolution of the convolution operation result and reduce the data dimension through average pooling operation, thereby improving the calculation efficiency and enhancing the robustness of the model. At the same time, by calculating the average value within the pooling window, the feature map is further smoothed to prevent overfitting and ensure that the model is more invariant to small changes in the input data.
[0184] The implementation of average pooling is specifically
[0185] Among them, represents the final pooling operation result, parameters such as represent the feature information extracted by the convolution kernel, represents the pooling process.
[0186] Perform average pooling on the convolution results of each channel and calculate the average value within the pooling window. For example
[0187] Among them, x i represents the th element within the pooling window, represents the total number of elements within the pooling window.
[0188] In this embodiment, an appropriate pooling window size (such as 2×2 or 3×3) is selected according to the spatial resolution of the feature map. A larger window can significantly reduce the dimension but may lose details; a smaller window retains more feature information.
[0189] The stride is usually the same as the pooling window size (for example, a 2×2 window uses a stride of 2) to avoid feature overlap. If more features are desired to be retained, a smaller stride (such as a stride of 1) can be selected.
[0190] In this embodiment, average pooling reduces the spatial resolution of the feature map, reduces the computational amount of subsequent network layers, and realizes dimension reduction and improvement of computational efficiency. Average pooling reduces the dependence on a single pixel point by taking the average value of the elements within the window, reduces the sensitivity of the model to noise, and prevents overfitting. The pooling process enhances the spatial invariance of the features, making the model more robust to small changes in the input data.
[0191] As a preferred embodiment of the present invention, the prediction of liquor consumption is obtained through linear rectification activation, specifically: Based on the pooling operation result, combined with the weight matrix and the bias vector, the prediction result of liquor consumption is obtained; Among them, the weight matrix is a two-dimensional matrix composed of the number of output neurons and the number of input neurons, and the bias vector is a one-dimensional array of the number of output neurons. The weight matrix and the bias vector are obtained through training.
[0192] This embodiment aims to globally integrate the feature map after convolution and pooling through a fully connected layer to extract higher-level abstract features. The introduction of the ReLU activation function further enhances the nonlinear expression ability of the model, enabling the network to better capture complex consumption behavior patterns.
[0193] The linear rectification activation formula is
[0194] Where is the output of the fully connected layer. is the input of the fully connected layer, containing the key feature information after dimensionality reduction of each channel. is the activation function, whose role is to introduce nonlinear characteristics to avoid the model being able to only fit linear relationships. Its characteristics include simple calculation, fewer gradient vanishing problems, and easy optimization. is the weight matrix, and the dimension of the weight matrix 𝑊 is determined by the size of the input feature map and the number of output neurons. The role of is to map the input features to a higher-dimensional or lower-dimensional space to meet the requirements of subsequent tasks (such as classification or regression). is the bias vector, which is used to adjust the offset of the output value and enhance the flexibility of the model.
[0195] This embodiment flattens the pooled feature map into a one-dimensional vector to ensure that all features can be uniformly processed by the fully connected layer. For example, if has a resolution of 4×4×9 (assuming a total of 9 channels), then a vector with a length of 144 is obtained after flattening. Use a suitable initialization method (such as Xavier initialization or He initialization) to set the weight matrix and the bias vector , to avoid the problems of gradient explosion or disappearance. Apply the ReLU activation function on the basis of weighted summation to enhance the nonlinear expression ability of the model, enabling it to capture more complex behavior patterns.
[0196] The fully connected layer of this embodiment globally integrates the local features of each channel to form higher-level abstract features, which are convenient for subsequent tasks to use. The introduction of the ReLU activation function enables the model to capture nonlinear relationships. By adjusting the weight matrix and the bias vector , the fully connected layer can flexibly adapt to different task objectives (such as classification or regression) and improve the versatility of the model.
[0197] As a preferred embodiment of this implementation manner, the weight matrix and the bias vector are obtained through training, specifically as follows: Determine the training data and the test data. Both the training data and the test data include the user's basic profile of liquor consumers, daily behavior trajectories, input data of purchase information results, and classification data obtained based on the purchase data of the current year; Train according to the input data and the classification data in the training data to obtain multiple groups of the weight matrix and the bias vector; According to the input data in the test data, combine multiple groups of the weight matrix and the bias vector to obtain multiple prediction results. Determine the weight matrix and the bias vector according to the correlation between the multiple prediction results and the classification data in the test data.
[0198] This embodiment aims to optimize the weight matrix and the bias vector so that the model can accurately predict user consumption classification; verify and select the optimal parameters through the test data to ensure the generalization ability of the model to new data.
[0199] It should be noted that in this embodiment, the data set is randomly divided according to a ratio (such as 80% training set and 20% test set) to ensure the same distribution of the two types of data. For example, a user with a "monthly average consumption of 5,000 yuan" is marked as a "high-value user" (classification data), and their behavior trajectory (such as "browsing red wine 3 times a week") is used as the input data. Ensure the distribution consistency of the test set and the training set through stratified sampling to avoid model bias.
[0200] Train the model according to the training data to obtain multiple groups of the weight matrix and the bias vector. First, initialize the parameters. For the weight matrix , use Xavier initialization to ensure a reasonable initial weight distribution and avoid gradient disappearance / explosion. For the bias vector , initialize it as a zero vector or a small random value to enhance the flexibility of the model.
[0201] During the training process, through forward propagation, the input data calculates the prediction result through the fully connected layer:
[0202] Among them, represents the th sample.
[0203] Perform loss calculation, and use the cross-entropy loss function to measure the difference between the prediction result and the classification data:
[0204] Among them, is the true label, is the predicted probability.
[0205] Back propagate again and calculate the loss pair by chain derivation and The gradient of . Through parameter update, the Adam optimizer is used to dynamically adjust the learning rate and update the parameters:
[0206] in, is the learning rate.
[0207] Generate multiple sets of parameters and generate multiple sets of (W, b) through multiple training (different initialization or hyperparameters) to avoid local optimality.
[0208] Select the optimal weight matrix and bias vector based on the test data. Input the test data into the model and use multiple sets of (W, b) to get the prediction results:
[0209] in, Indicates the Group parameters.
[0210] Select the (W, b) that performs best on the test set. Note that if a set of parameters has a low training loss but a high test loss, it should be discarded to prevent overfitting.
[0211] The present invention also provides an electronic device, comprising: Memory, for storing computer instructions; A processor is used to implement the wine consumption prediction method combining curvelet transform and convolutional neural network when executing the computer instructions.
[0212] Therefore, this electronic device can achieve any effect of the wine consumption prediction method combining curvelet transform and convolutional neural network, which will not be described in detail here.
[0213] Anything not described in the present invention can be achieved by adopting or drawing on existing technologies.
[0214] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0215] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A liquor consumption prediction method combining curvelet transform and convolutional neural network, characterized in that, Including: Obtain a time - feature data set based on the user basic profile, daily behavior trajectory, and purchase information results of liquor - consuming users; According to the time - feature data set, through curvelet transform and combined with curvelet characteristic frequencies, obtain multiple curvelet coefficients, and the multiple curvelet coefficients include low - frequency curvelet coefficients, medium - frequency curvelet coefficients, high - frequency curvelet coefficients, and ultra - high - frequency curvelet coefficients; According to the low - frequency curvelet coefficients, medium - frequency curvelet coefficients, and high - frequency curvelet coefficients, through eigen - decomposition, obtain multiple feature information; According to the multiple feature information, through a convolutional neural network, obtain multiple convolution operation results, obtain a pooling operation result through pooling operation, and obtain liquor consumption prediction through linear rectification activation.
2. The method for predicting liquor consumption combining curvelet transform and convolutional neural network according to claim 1, characterized in that, Obtain a time - feature data set according to the user basic profile, daily behavior trajectory, and purchase information results of liquor - consuming users, specifically: Perform formatted conversion according to the user basic profile, daily behavior trajectory, and purchase information results of liquor - consuming users, and obtain time - series numerical values in one - to - one correspondence; According to the time - series numerical values, obtain a time - feature data set.
3. The method for predicting liquor consumption combining curvelet transform and convolutional neural network according to claim 2, characterized in that, The user basic profile, daily behavior trajectory, and purchase information results of liquor - consuming users are specifically: The user basic profile includes at least any one of the consumer's age, gender, date of birth, and registration time; The daily behavior trajectory includes at least any one of the time of logging in to the platform, the time of browsing products, the time of pulling up payment, the number of order cancellations, the number of favorited products, and the number of activity participations; The purchase result information includes at least any one of the purchase amount, quantity, purchase time, and purchase store.
4. The method for predicting liquor consumption combining curvelet transform and convolutional neural network according to claim 1, characterized in that, According to the time - feature data set, through curvelet transform and combined with curvelet characteristic frequencies, obtain multiple curvelet coefficients, specifically: According to the time - feature data set, through curvelet transform, obtain curvelet coefficients at 6 scales; Combined with curvelet characteristic frequencies, divide the curvelet coefficients at 6 scales into low - frequency curvelet coefficients, medium - frequency curvelet coefficients, high - frequency curvelet coefficients, and ultra - high - frequency curvelet coefficients; Among them, both the medium - frequency curvelet coefficients and the high - frequency curvelet coefficients contain curvelet coefficients at 2 scales.
5. The method for predicting liquor consumption combining curvelet transform and convolutional neural network according to claim 4, characterized in that, According to the low - frequency curvelet coefficients, medium - frequency curvelet coefficients, and high - frequency curvelet coefficients, through eigen - decomposition, obtain multiple feature information, specifically: According to the low - frequency curvelet coefficients, through inverse curvelet transform, obtain trend feature information; According to the medium - frequency curvelet coefficients, through multiple - direction classification in the medium - frequency curvelet coefficients and through inverse curvelet transform, obtain medium - frequency feature information in 4 directions; According to the high - frequency curvelet coefficients, through multiple - direction classification in the high - frequency curvelet coefficients and through inverse curvelet transform, obtain high - frequency feature information in 4 directions.
6. The method for predicting liquor consumption combining curvelet transform and convolutional neural network according to claim 5, wherein According to the multiple feature information, through a convolutional neural network, obtain multiple convolution operation results, specifically: According to the trend feature information, through a sharpening convolution kernel, obtain a convolution operation result; According to the horizontal medium - frequency feature information and high - frequency feature information, through a vertical edge - detection convolution kernel, extract the longitudinal gradient change, and respectively obtain convolution operation results; According to the longitudinal medium - frequency feature information and high - frequency feature information, through a horizontal edge - detection convolution kernel, extract the horizontal gradient change, and respectively obtain convolution operation results; According to the oblique intermediate frequency feature information, a convolution operation result consistent with the oblique intermediate frequency feature information is obtained through a unit convolution kernel; According to the oblique high-frequency feature information, a convolution operation result is obtained through a convolution kernel for image smoothing.
7. The method for predicting liquor consumption by combining curvelet transform and convolutional neural network according to claim 1, characterized in that A pooling operation result is obtained through a pooling operation, specifically: According to multiple convolution operation results, an average pooling operation result representing the average value within the pooling window is obtained through average pooling to reduce the data dimension.
8. The method for predicting liquor consumption by combining curvelet transform and convolutional neural network according to claim 1, wherein The prediction of liquor consumption is obtained through linear rectification activation, specifically: According to the pooling operation result, in combination with a weight matrix and a bias vector, a liquor consumption prediction result is obtained; Wherein the weight matrix is a two-dimensional matrix composed of the number of output neurons and the number of input neurons, and the bias vector is a one-dimensional array of the number of output neurons, and the weight matrix and the bias vector are obtained through training.
9. The method for predicting liquor consumption combining curvelet transform and convolutional neural network according to claim 8, characterized in that, The weight matrix and the bias vector are obtained through training, specifically: Determine training data and test data, and both the training data and the test data include the user basic profile of liquor consumption users, daily behavior trajectories, input data of purchase information results, and classification data obtained according to the purchase data of the current year; Train according to the input data and classification data in the training data to obtain multiple groups of the weight matrix and the bias vector; According to the input data in the test data, in combination with multiple groups of the weight matrix and the bias vector, multiple prediction results are obtained, and according to the correlation between the multiple prediction results and the classification data in the test data, the weight matrix and the bias vector are determined.
10. An electronic device, characterized in that, Including: A memory for storing computer instructions; A processor for implementing the liquor consumption prediction method combining curvelet transform and convolutional neural network according to any one of claims 1 to 9 when executing the computer instructions.
Citation Information
Patent Citations
A Short-Term Power Load Forecasting Method
CN102270279A
Convolutional neural network face recognition algorithm based on wavelet transform and DCT
CN110598584A
Power supply load prediction method based on convolutional neural network
CN115358437A
Cigarette sales volume prediction method and device, computer equipment and storage medium
CN116701944A
White spirit quality prediction method
CN116805180A