Amazon commodity score star level prediction method based on long and short term memory neural network forest
Through the method based on long and short-term memory neural network forest, the feature data of Amazon products are extracted and modeled, the gap in Amazon product star rating prediction is solved, accurate prediction of future star ratings is achieved, and more effective inventory and operation management is supported.
Patent Information
- Application Number
- CN202411309052.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-06-03
AI Technical Summary
There is a lack of effective methods for predicting Amazon product star ratings in the prior art, and it is difficult to accurately predict the highest and lowest star ratings of monthly pages for a period of time in the future.
Using a method based on long and short-term memory neural network forest, we collect and process Amazon product details and comment data, extract feature data and build a long and short-term memory neural network forest model to perform star-rated predictions.
Accurate predictions of the highest and lowest star ratings of Amazon products in the next 6 months have been achieved, helping sellers better manage inventory and operation strategies.
Smart Images

Figure CN120087997A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural networks, and particularly relates to a method for predicting the Amazon product rating star based on a long short-term memory neural network forest. Background Art
[0002] The Amazon star rating is comprehensively affected by product quality, product price, product selling points, etc. The product star rating on Amazon, especially the highest star rating and the lowest star rating on the monthly page, will directly affect the sales volume of the product. If the changes in the highest star rating and the lowest star rating on the monthly page in the future period can be known, it can provide reference value for inventory management, replenishment planning, operation strategies, etc.
[0003] In the prior art, the long short-term memory neural network forest is only an algorithm idea. For specific problems, relevant features of the problem need to be determined, the time window and granularity of the features need to be determined, and a long short-term memory neural network forest model structure suitable for the specific problem needs to be constructed. Therefore, for the problem of predicting the star rating of Amazon products, there is currently no corresponding technology to handle it. Summary of the Invention
[0004] The present invention discloses a method for predicting the Amazon product rating star based on a long short-term memory neural network forest. The method for predicting the Amazon product rating star based on a long short-term memory neural network forest includes the following steps:
[0005] Step 1, collect data of the product details page and data of the product review page of the product to be predicted on the Amazon platform;
[0006] Step 2, extract the web page label content from the data of the product details page and the data of the product review page through document parsing and store it in the database;
[0007] Step 3, extract feature data for constructing the long short-term memory neural network forest model according to the web page label content;
[0008] Step 4, construct a long short-term memory neural network forest model for prediction;
[0009] Step 5, train and store the neural network forest models of each product to be predicted respectively;
[0010] Step 6, use the data of each product to be predicted recently to input into the neural network forest model for rating star prediction;
[0011] Step 7, visually display the prediction results in the form of a chart.
[0012] Further, in step 1, the data of the product details page includes the average page star rating;
[0013] The data of the product review page includes: review star rating 1, review time 2, cumulative number of reviews, and cumulative number of reviews in the previous month.
[0014] Further, in step 2, the web page tag content includes the average page star rating, review star rating 1, review time 2, cumulative number of reviews, and cumulative number of reviews in the previous month.
[0015] Further, in step 3, the feature data includes the number of newly added reviews per month, the average star rating of newly added reviews per month, the cumulative number of reviews per month, the total cumulative number of reviews per month, the average page star rating per month, the average page star rating, the annual dimension, and the monthly dimension;
[0016] The highest star rating per month on the page and the lowest star rating per month on the page are used as the prediction targets.
[0017] Further, in step 3, the number of newly added reviews per month is the total number of reviews in the current month minus the total number of reviews in the previous month;
[0018] The average star rating of newly added reviews per month is the average star rating of reviews in the current month minus the average star rating of reviews in the previous month;
[0019] The cumulative review data per month is the sum of the number of reviews in the current month;
[0020] The total cumulative number of reviews per month is the sum of the number of reviews on this page as of the current month;
[0021] The average page star rating per month is the sum of the star ratings corresponding to each review in the current month divided by the total number of days in the current month;
[0022] The average page star rating is the weighted average of the star ratings corresponding to all reviews of this product;
[0023] The highest star rating per month on the page is the largest value among the star ratings corresponding to the reviews in the current month, and the maximum value of this value is 5;
[0024] The lowest star rating per month on the page is the smallest value among the star ratings corresponding to the reviews in the current month, and the minimum value of this value is 1;
[0025] The annual dimension is the year data corresponding to the current month; the monthly dimension is the month data of the current month, and the value range is from 1 to 12.
[0026] Further, in step 4, the following steps are also included:
[0027] Step 41, constructing a feature cluster through the feature data;
[0028] Step 42, inputting the feature cluster into the long short-term memory neural network forest model.
[0029] Further, in step 41, all x assignable features of the data are obtained, m feature numbers are randomly selected without replacement from 1 - x to form a complete feature cluster, and this process is repeated n times until n feature clusters are selected. Among them, the year data and month data are non - assignable features and must form new feature sets with each feature cluster.
[0030] Further,
[0031] In step 42, the long - short - term memory neural network forest model includes multiple long - short - term memory neural networks and a fully - connected layer. Each long - short - term memory neural network includes an input layer, a hidden layer, and an output layer. The outputs of multiple long - short - term memory neural networks are processed through the fully - connected layer to achieve the processing of the output results of all long - short - term memory neural networks;
[0032] The input layer is used to form multiple feature clusters;
[0033] The hidden layer includes 2 layers, each layer has 10 neurons, the forgetting coefficient is 0.2, the activation function is selected as the RELU function, and the learning rate is 0.0002;
[0034] The time window of the output layer is 6, the number of target predictions is 2, and the size of the output layer is 6 * 2;
[0035] The fully - connected layer performs weighted calculations on the output results of all LSTMs and obtains the prediction results. The time window of the fully - connected layer is 6, the number of target predictions is 2, and the size of the output layer is 6 * 2. Further, in step 5, the long - short - term memory neural network forest models for predicting the star ratings of each product page are stored separately. The long - short - term memory neural network forest models are saved in h5 format and stored on the disk in file form.
[0036] Further, in step 5, during the training phase, 80% of the historical feature data of each product is used as the training data set, and 20% is used as the test data set;
[0037] At the same time, a misaligned training method is adopted, that is, the feature values and target values are misaligned and mapped, and 6 - month historical data is reserved for predicting the highest star rating and the lowest star rating of the product page in the next 6 months.
[0038] The beneficial effects achieved by the present invention are:
[0039] The present invention innovatively draws on the idea of the random forest composition, uses multiple long - short - term memory neural networks to form a long - short - term memory neural network forest, and applies this algorithm to the prediction of Amazon product rating star levels, filling the gap in the prediction of Amazon product rating star levels. It helps Amazon sellers better manage product sales, inventory management, and product iteration. Description of the Drawings
[0040] Figure 1 Schematic flowchart of a method for predicting the Amazon product rating star based on a long short-term memory neural network forest provided by the present invention;
[0041] Figure 2 Schematic diagram of the product page rating star in a method for predicting the Amazon product rating star based on a long short-term memory neural network forest provided by the present invention;
[0042] Figure 3 Schematic diagram of the product page comments in a method for predicting the Amazon product rating star based on a long short-term memory neural network forest provided by the present invention;
[0043] Figure 4 Schematic diagram of the structure of the neural network model in a method for predicting the Amazon product rating star based on a long short-term memory neural network forest provided by the present invention;
[0044] Figure 5 Relationship diagram showing that the error values gradually decrease with the increase of the iteration number when the neural network model provided by the present invention predicts the star ratings of two different products during the training process in a method for predicting the Amazon product rating star based on a long short-term memory neural network forest;
[0045] Figure 6 Trend chart of the product star rating prediction by a method for predicting the Amazon product rating star based on a long short-term memory neural network forest provided by the present invention;
[0046] Figure 7 Trend chart of the product star rating prediction by a method for predicting the Amazon product rating star based on a long short-term memory neural network forest provided by the present invention. Detailed implementation manner
[0047] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, these embodiments are exemplary only and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that the details and forms of the technical solutions of the present invention can be modified or replaced without departing from the spirit and scope of the present invention, but such modifications and replacements all fall within the protection scope of the present invention.
[0048] The object of the present invention is to predict the product star rating for the next 6 months.
[0049] Such as Figure 1 The main processes of the present invention involve data collection, data processing, feature extraction, model construction, model training and storage, model application, prediction result visualization, and workflow construction.
[0050] Step 1, Data Collection: Use web crawler technology to collect data from the product detail pages and product review pages on Amazon.
[0051] The data on the product detail page includes the average star rating of the page;
[0052] The data on the product review page includes: review star rating 1, review time 2, cumulative number of reviews, and cumulative number of reviews in the previous month.
[0053] Step 2, Data Processing: Based on the data collected from the product detail page and the product review page according to the above data collection process, extract the web page tag content through document parsing and store it in the database.
[0054] The web page tag content includes the average star rating of the page, review star rating 1, review time 2, cumulative number of reviews, and cumulative number of reviews in the previous month.
[0055] Step 3, Feature Extraction: According to the result data in the database after the above data processing, extract the review star rating 1, review time 2, cumulative number of reviews, cumulative number of reviews in the previous month, etc. on the Amazon review page as shown in Figure 3 , and summarize and statistically analyze the above feature values monthly to obtain the number of new reviews per month, the average star rating of new reviews per month, the cumulative number of reviews per month, the total cumulative number of reviews per month, the average star rating of the page per month, the highest star rating of the page per month, and the lowest star rating of the page per month. Extract the average star rating (average star rating of the page) on the Amazon product detail page as shown in Figure 2 and its collection time. At the same time, combine the year and month dimensions of time for a total of 10 quantities. The highest star rating of the page per month and the lowest star rating of the page per month are used as labels, and the remaining 8 quantities are used as features.
[0056] Among them, the number of new reviews per month is the total number of reviews in the current month minus the total number of reviews in the previous month; the average star rating of new reviews per month is the average of the star ratings of reviews in the current month minus the average of the star ratings of reviews in the previous month; the cumulative review data per month is the sum of the number of reviews in the current month; the total cumulative number of reviews per month is the sum of the number of reviews of this product as of the current month; the average star rating of the page per month is the sum of the star ratings corresponding to each review in the current month divided by the total number of days in the current month; the average star rating of the page is calculated as the weighted average of the star ratings corresponding to all reviews of this product; the year dimension is the year data corresponding to the current month; the month dimension is the month data of the current month, with a value range of 1 to 12; the highest star rating of the page per month is the largest value among the star ratings corresponding to the reviews in the current month, and the maximum value of this value is 5; the lowest star rating of the page per month is the smallest value among the star ratings corresponding to the reviews in the current month, and the minimum value of this value is 1.
[0057] Finally, normalize the above 8 feature data.
[0058] Step 4, Model construction:
[0059] As Figure 4 shown, generally speaking, first, construct n non-interfering feature clusters by randomly selecting features n times. Secondly, input the data values related to each feature cluster into the corresponding long short-term memory neural network respectively to predict the relevant information carried by each feature cluster about six months later. Finally, perform a weighted average on all the information through a fully connected layer to obtain the final predicted value.
[0060] Step 41, Construction of feature clusters:
[0061] First, obtain all x assignable features of the data, randomly select m feature numbers from 1 - x without replacement to form a complete feature cluster, and repeat this process n times until n feature clusters are selected. Among them, the year data and month data are non-assignable features and must form new feature sets with each feature cluster.
[0062] In this patent, there are 2 non-assignable features and 6 assignable features. Randomly select 3 features from all assignable features (for example, the newly added comments per month, the cumulative comments per month, and the highest star rating of the page per month) to form feature cluster A, and form feature cluster A+ with the two non-assignable features. Therefore, input the data of the first six months under these 5 features into the long short-term memory neural network under A+ for learning to obtain the prediction situation of the network under feature cluster A+.
[0063] A random forest consists of multiple weak classifier decision forests. When using a random forest, the number of decision forests and the number of features of the feature clusters need to be set. Therefore, in this patent, the number of feature clusters and long short-term memory neural networks and the number of features carried by the feature clusters also need to be set.
[0064] Step 42, Input the feature clusters into each long short-term memory neural network forest;
[0065] In the present invention, the number of features carried by the feature clusters is 5, and the total number of feature clusters constructed through Step 41 is That is, the number of feature clusters and long short-term memory neural networks is 20;
[0066] The long short-term memory neural network forest model therein contains multiple long short-term memory neural networks. Each long short-term memory neural network includes an input layer, a hidden layer, and an output layer. The outputs of multiple long short-term memory neural networks are processed through a fully connected layer as the total output layer to realize the processing of the output results of all long short-term memory neural networks. The model structure is as Figure 4 shown;
[0067] Input layer: After step 41, multiple feature clusters are formed. In the present invention, the number of features in each feature cluster is 5, and the time window is 6. Therefore, the size of the input layer is 6 * 5.
[0068] Hidden layer: There are a total of 2 layers, with 10 neurons in each layer. The forgetting coefficient is 0.2, the activation function is the RELU function, and the learning rate is 0.0002.
[0069] Output layer: The last layer of the LSTM model. The time window is 6, and the number of target predictions is 2. So the size of the output layer of the LSTM is 6 * 2.
[0070] Fully connected layer: The role of the fully connected layer is to perform weighted calculations on the output results of the above 20 LSTMs and obtain the prediction results. The time window is 6, and the number of target predictions is 2. So the size of the fully connected layer is also 6 * 2.
[0071] The W weight adjustment algorithm uses the ADAM algorithm.
[0072] The loss function is the mean squared error.
[0073] Step 5, model training and storage:
[0074] Since the star ratings of different Amazon products are independent, the star rating of a product mainly reflects the characteristics of the product itself and its historical situation. Therefore, in the present invention, for each product, as many models as there are products are output for star rating prediction. At the same time, the models are saved in the h5 format and stored on the disk in the form of files.
[0075] In the training stage, 80% of the historical feature data of each product is used as the training data set, and 20% is used as the test data set. For the convenience of model application, this patent adopts the method of misaligned training, that is, the feature values and target values are misaligned and mapped, and 6 months of historical data is reserved for predicting the highest star rating and the lowest star rating of the product page in the next 6 months. This patent uses the ADAM algorithm to update the above various W weight values. Then, the trained model is input with the test data set to judge the mean squared error value of the prediction. Figure 5 It is a relationship graph showing that the error value gradually decreases as the number of iterations increases when predicting the star ratings of two different products during the training process.
[0076] Step 6, using the constructed model to predict future data:
[0077] Load the model corresponding to the trained h5 format file, input the feature cluster data described in the above feature extraction steps of the product in the recent 6 months, predict the highest star rating and the lowest star rating of the product page in the next 6 months, and store the above prediction results in the database.
[0078] Step 7, Visualization of prediction results: The historical data and prediction data saved in the database during the above model application phase are presented in the form of charts through BI software, such as Figure 6 , Figure 7 . Figure 6 The product model in Figure 7 was trained and generated on September 1, 2022. The highest star rating and the lowest star rating on the page are the prediction results. Therefore, the data before September 2022 belong to historical data, and the data from September 2022 to February 2023 belong to prediction data. From the historical data, it can be found that as the shelf time of the product increases and the number of comments increases, the changes in the highest and lowest star ratings on the page tend to be stable. From April 2022 to August 2022, the highest and lowest star ratings on the page are basically the same. However, when making predictions, the LSTM algorithm will take into account the error factors. Therefore, the highest predicted star rating is slightly higher than the recent page star rating data, and the lowest value is slightly lower than the page star rating data. Figure 6 is the prediction of the highest and lowest star ratings on the page of another product. Similar to the above
[0079] The data from September 2022 to February 2023 is predicted. The algorithm accurately takes into account that the average star rating of the new comments in the recent six months is lower than the highest and lowest star ratings on the product page. Therefore, there is a downward trend in the highest and lowest star ratings on the page in the next six months. After analyzing all the remaining products, it is found that the overall effect meets the expected requirements.
[0080] The present invention is not limited to the above specific embodiments. Those of ordinary skill in the art can implement the present invention in various other specific embodiments according to the embodiments and the disclosed content of the drawings. Therefore, any design that adopts the design structure and concept of the present invention and makes some simple transformations or changes falls within the protection scope of the present invention.
Claims
1. A method for predicting Amazon product ratings based on long short-term memory neural network forests, characterized in that: The Amazon product rating star prediction method based on long short-term memory neural network forest includes the following steps: Step 1: Collect data on the product details page and product review page to be predicted on the Amazon platform; Step 2: extract web page tag content from the product details page and product review page through document parsing and store them in the database; Step 3, extracting feature data for constructing a long short-term memory neural network forest model based on the web page tag content; Step 4, construct a long short-term memory neural network forest model for prediction; Step 5, training and storing the neural network forest model of each commodity to be predicted; Step 6: Use the recent data of the products to be predicted to input into the neural network forest model to predict the star rating; Step 7: Visualize the prediction results in the form of charts.
2. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 1, characterized in that: In step 1, the data of the product details page includes the average star rating of the page; The data on the product review page includes: review star (1), review time (2), cumulative number of reviews, and cumulative number of reviews last month.
3. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 1, characterized in that: In step 2, the web page tag content includes the average star rating of the page, the star rating of the review (1), the review time (2), the cumulative number of reviews, and the cumulative number of reviews in the previous month.
4. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 3, characterized in that: In step 3, the characteristic data includes the number of new comments per month, the star rating of new comments per month, the cumulative number of comments per month, the total cumulative number of comments per month, the average star rating of pages per month, the average star rating of pages, the year dimension, and the month dimension; The highest page star rating per month and the lowest page star rating per month are used as prediction targets.
5. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 4, characterized in that: In step 3, the number of new comments each month is the total number of comments in that month minus the total number of comments in the previous month; The star rating of newly added reviews each month is the average star rating of reviews in that month minus the average star rating of reviews in the previous month; Monthly cumulative review data is the sum of the number of reviews in that month; The total number of cumulative comments each month is the sum of the number of comments on the page as of that month; The average star rating of the page each month is the sum of the star ratings of each review in that month, divided by the total number of days in that month; The average star rating of the page, which is the weighted average of the star ratings of all reviews of the product; The highest star rating of each month’s page is the highest star rating among the corresponding star ratings of the reviews of that month, and the maximum value is 5; The lowest star rating of each month’s page is the lowest star rating among the corresponding star ratings of the reviews of that month, and the minimum star rating is 1. The year dimension refers to the year data corresponding to the current month; the month dimension refers to the month data of the current month, and the value range is 1 to 12.
6. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 1, characterized in that: In step 4, the following steps are also included: Step 41, constructing a feature cluster through feature data; Step 42, input the feature cluster into the long short-term memory neural network forest model.
7. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 6, characterized in that: In step 41, all x assignable features of the data are obtained, and m features are randomly selected from 1-x without replacement to form a complete feature cluster. This process is repeated n times until n feature clusters are selected. Among them, the year data and the month data are non-assignable features and must form a new feature set with each feature cluster.
8. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 6, characterized in that: In step 42, the LSTM neural network forest model includes multiple LSTM neural networks and fully connected layers, each LSTM neural network includes an input layer, a hidden layer and an output layer, and the outputs of the multiple LSTM neural networks are processed through the fully connected layer to realize the output results of all LSTM neural networks; The input layer is used to form multiple feature clusters; The hidden layer consists of 2 layers, each with 10 neurons, the forgetting coefficient is 0.2, the activation function uses the RELU function, and the learning rate is 0.0002; The time window of the output layer is 6, the number of target predictions is 2, and the size of the output layer is 6*2; The fully connected layer performs weighted calculation on the output results of all LSTMs and obtains the prediction results. The time window of the fully connected layer is 6, the number of target predictions is 2, and the size of the output layer is 6*2.
9. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 1, characterized in that: In step 5, the long short-term memory neural network forest model for star rating prediction of each product page is stored separately, and the long short-term memory neural network forest model is saved in h5 format and stored on the disk as a file.
10. The Amazon product rating star prediction method based on long short-term memory neural network forest according to claim 1, characterized in that: In step 5, 80% of the historical feature data of each product in the training phase is used as a training data set, and 20% is used as a test data set; At the same time, the staggered training method is adopted, that is, the feature value and the target value are staggered and mapped, and 6 months of historical data are reserved for the prediction of the highest star rating and the lowest star rating of the product page in the next 6 months.