Business order analysis and prediction system based on big data
By designing a business order analysis and prediction system based on big data, using a variety of machine learning algorithms and text analysis technologies, the limitations of existing business analysis tools when processing large-scale data sets are solved, and accurate classification and trend prediction of business orders are achieved, and resource utilization efficiency is improved.
Patent Information
- Application Number
- CN202411312071.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-09-20
AI Technical Summary
Existing business analysis tools have limitations when dealing with large-scale data sets, requiring analysis of different data types than users want to analyze, resulting in waste of resources.
Design a business order analysis and prediction system based on big data, including classification module, determination module and analysis module. The classification module classifies business orders through text analysis methods, uses multiple machine learning algorithms to train classification models, and optimizes model performance. The determination module determines the business order type based on user needs, and the analysis module conducts trend analysis and prediction.
Through multi-model fusion and text analysis technology, the system can accurately extract key features of business orders, improve the accuracy of classification and trend prediction, avoid the analysis of irrelevant data, optimize resource allocation, and improve data processing efficiency.
Smart Images

Figure CN118820985B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a business order analysis and prediction system based on big data. Background Art
[0002] With the development of big data technology, enterprises are able to collect and process a large amount of business order data, but how to extract valuable information from this data, predict future trends, and formulate strategies accordingly is still a challenge. Existing business analysis tools have limitations when processing large-scale data sets, and often need to analyze data types in big data that are different from the data that users want to analyze, resulting in a waste of resources. Therefore, a business order analysis and prediction system based on big data is designed.
[0003] In order to solve the above defects, a technical solution is now provided. Summary of the invention
[0004] The purpose of the present invention is to solve the limitation of existing business analysis tools in processing large-scale data sets, which often requires analyzing data types in big data that are different from those that users want to analyze, resulting in a waste of resources, and propose a business order analysis and prediction system based on big data.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] The business order analysis and prediction system based on big data includes:
[0007] The classification module is used to classify business orders based on text analysis methods. The specific steps are as follows:
[0008] First, obtain the business order text information, perform data preprocessing on the business order text information to remove noise and standardize the data format;
[0009] Then extract keywords or phrases as features from the preprocessed business order text information:
[0010] Word frequency statistics: count the frequency of each word or phrase appearing in the text;
[0011] Term Frequency-Inverse Document Frequency: Use the term frequency-inverse document frequency method to evaluate the importance of a word to a document in a text collection;
[0012] N-gram model: extracts combinations of N consecutive words to better capture word order information;
[0013] Then, three machine learning algorithms, Naive Bayes, Support Vector Machine, and Deep Learning, were used to train the classification model;
[0014] The trained model is then applied to the preprocessed business order text information, and the orders are classified according to the keywords in the business order text information: by setting a threshold, when the model's predicted probability exceeds this threshold, the order is classified into the corresponding category; when the order belongs to multiple categories, a multi-label classification method is used;
[0015] Then evaluate the performance of the classification model and optimize it based on the feedback;
[0016] The determination module is used to determine the type of business order based on the user's analysis requirements;
[0017] The analysis module is used to analyze the business order trends that users need to query.
[0018] Furthermore, the specific process of the classification module evaluating the classification model is as follows:
[0019] The classification model is evaluated by obtaining the accuracy, recall, precision and F1 score of the classification model after classification.
[0020] The accuracy rate is a measure of the proportion of orders correctly classified by the model. It is a model performance evaluation indicator, indicating the proportion of orders correctly classified by the model to the total number of orders. The calculation formula is: , Z is the accuracy, q and s are the number of correctly classified orders and the total number of orders respectively;
[0021] The recall rate measures the ability of the model to correctly identify all orders. It measures the proportion of positive classes (related orders) correctly identified by the model to all actual positive classes. The calculation formula is: , where H is the recall rate, tp and fn are true positive examples and false negative examples respectively;
[0022] Precision measures the proportion of orders that are actually positive among those identified as positive by the model. The calculation formula is: , where J is the precision and fp is the false positive example;
[0023] The F1 score is an indicator that reconciles precision and recall. It is used to measure the overall performance of the model. The calculation formula is: , where F1 is the F1 score. The larger the calculated F1 score is, the better the performance of the corresponding classification model is.
[0024] Then normalize the obtained accuracy, recall, precision and F1 score and enter them into the following formula: To get the comprehensive evaluation value ZPZ, where They are the preset weight coefficients for accuracy, recall, precision and F1 score respectively;
[0025] The calculated comprehensive evaluation value ZPZ is then compared with the preset comprehensive evaluation threshold. When the comprehensive evaluation value is lower than the preset comprehensive evaluation threshold, it is judged that the classification model effect does not meet the standard, and the model is optimized based on the feedback.
[0026] Furthermore, the specific operation steps of the classification module to optimize the model according to the feedback are as follows:
[0027] Feedback sources include:
[0028] User feedback: Gather feedback from business teams and end users who use the classification model on model prediction accuracy, user experience, and actual business impact;
[0029] Business Indicators: Monitor business indicators, including customer satisfaction, order processing time, and inventory turnover;
[0030] Error analysis: Analyze cases where the model made incorrect predictions;
[0031] Market changes: Pay attention to changes in market trends and customer needs;
[0032] The specific optimization items are as follows:
[0033] Adjust model parameters: Based on performance indicators and user feedback, adjust the model's hyperparameters, including learning rate, tree depth, and regularization strength;
[0034] Feature Engineering: Improve the feature extraction and selection process based on error analysis and user feedback, including adding new features, modifying feature conversion methods, or removing irrelevant features;
[0035] Model selection and integration: When it is determined that a single model cannot meet the performance requirements, use alternative classification models or adopt ensemble learning methods, including random forests, gradient boosting machines, or ensembles of neural networks;
[0036] Dataset update and expansion: Update and expand the training dataset based on market changes and newly collected user feedback, including adding new samples, rebalancing category distribution, or introducing new data sources;
[0037] Continuous Monitoring and Evaluation: Establish an ongoing monitoring and evaluation mechanism to regularly check the performance of the classification model and make adjustments based on the latest data and feedback.
[0038] Furthermore, the specific operation steps of the determination module for determining the type of the user's analysis demand service order are as follows:
[0039] After completing the classification of the business orders, the user's business order query requirements are obtained, and key features are extracted according to the user's query requirements, and corresponding features are extracted from each business order that has been classified;
[0040] Then, the user demand features and each classified business order features are converted into numerical feature vectors through word frequency, word frequency-inverse document frequency or other text vectorization methods, and the user demand features and the classified business order features are marked as A and B respectively;
[0041] Then three similarity calculation methods are used to calculate the similarity between the user demand feature vector and each classified business order feature vector;
[0042] Then, the above three similarity calculation methods are comprehensively analyzed, specifically:
[0043] Each similarity score is standardized to the range of [0, 1]. Then, different weights are assigned to each similarity method according to business needs or the importance of features. The final comprehensive similarity score is calculated using weighted average or simple average. The weights of the three similarity calculation methods are set to w1, w2, and w3, respectively, and the scores are S1, S2, and S3, respectively. The comprehensive similarity score is calculated using the following formula: ;
[0044] The similarity between the characteristics of user needs and the characteristics of different classified business orders is calculated respectively to obtain different comprehensive similarity scores, and the different comprehensive similarity scores are sorted according to size. The largest comprehensive similarity score is selected, and the corresponding classified business order type is used as the type of user business order query demand.
[0045] Furthermore, the specific operation steps of the determination module using three similarity calculation methods to calculate the similarity between the user demand feature vector and each classified business order feature vector are as follows:
[0046] Similarity calculation methods include:
[0047] Cosine similarity: For vectors A and B, the similarity between them is determined by measuring the cosine value of the angle between the two vectors. The formula is: , A·B is the dot product of the vectors, ‖A‖ and ‖B‖ are the Euclidean norms or lengths of the vectors respectively;
[0048] Euclidean distance; for vectors A and B, the Euclidean distance is the straight-line distance between A and B, and the formula is: , when the Euclidean distance is smaller, the similarity is higher;
[0049] Jaccard similarity: For two sets A and B, the Jaccard similarity is the size of the intersection of the two sets divided by the size of the union. The formula is: , |A ∩ B| is the number of elements in the intersection of sets A and B, and |A ∪ B| is the number of elements in the union of sets A and B.
[0050] Furthermore, the specific operation steps of the analysis module for analyzing the business order trend that the user needs to query are as follows:
[0051] First, the data of the business order types identified in the big data are preprocessed, including data cleaning, integration, normalization and processing of missing values; statistical tests, correlation analysis and feature selection algorithms are used to identify features with predictive value on the preprocessed data, thereby determining features that have an impact on trend prediction;
[0052] Then establish a time series analysis model to predict future trends. The time series models include autoregression, moving average, autoregression moving average and autoregression integrated sliding average models, and use random forest, support vector machine, gradient boosting machine and neural network machine learning algorithms to predict trends.
[0053] Then the model is trained and validated. Specifically, the selected model is trained using historical order data, and the prediction performance of the model is verified by cross-validation or by reserving a portion of the data as a test set. The evaluation indicators include mean square error, root mean square error, and mean absolute error.
[0054] Use the trained model to predict future order trends and predict short-term or long-term trends based on business needs;
[0055] Then, a correction model is established based on the prediction results of the model, and the prediction results are corrected through the correction model.
[0056] Furthermore, the specific operation steps of the analysis module to correct the prediction results by correcting the model are as follows:
[0057] First, identify and collect key features that affect order volume, including seasonal factors, promotional activities, economic indicators, competitor behavior, market trends, changes in customer behavior, and product updates or inventory levels;
[0058] Then, we use crawler technology, application programming interface calls and third-party data services to collect relevant data on the above key features and conduct preliminary analysis to obtain the relationship between these features and order volume;
[0059] Then perform feature engineering on the collected data, including feature extraction, transformation and selection, using time series analysis methods to extract seasonal components, or using statistical methods to determine which features are most correlated with order volume;
[0060] Establish a correction model, adjust the output of the basic prediction model according to the changes in key features through the correction model, and then use the same type of business order data and key feature data to train the correction model;
[0061] Finally, the output of the basic prediction model is used as input, combined with key feature data, and a correction model is used to adjust the prediction results.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] (1) The present invention uses a variety of machine learning algorithms and text analysis techniques to extract key features from a large amount of business order data and build accurate classification models and trend prediction models based on them. This multi-model fusion method helps to improve the prediction accuracy of business order trends;
[0064] (2) The present invention accurately identifies and screens the types of business orders that users are interested in through user demand analysis and similarity calculation methods, thereby avoiding invalid analysis of irrelevant data types, optimizing resource allocation, and improving the efficiency and practicality of data processing;
[0065] (3) The present invention designs a correction model that can dynamically correct the results of the basic prediction model based on the key features collected in real time. The real-time adjustment mechanism enables the system to adapt to market changes and data drift, maintaining the timeliness and flexibility of the prediction results. At the same time, the system also includes a continuous monitoring and evaluation mechanism to ensure the continuous optimization and adjustment of the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to facilitate understanding by those skilled in the art, the present invention is further described below in conjunction with the accompanying drawings;
[0067] Figure 1 This is the overall system block diagram of the present invention. DETAILED DESCRIPTION
[0068] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0069] It should be understood that the terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0070] It should also be understood that the terms used in this disclosure specification are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure specification and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" used in this disclosure specification and the claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0071] As Figure 1 shown, the business order analysis and prediction system based on big data includes a classification module, a determination module, and an analysis module;
[0072] The classification module is used to classify business orders according to the text analysis method. The specific steps are as follows:
[0073] First, obtain the business order text information and perform data preprocessing on the business order text information to eliminate noise and standardize the data format, including: removing stop words: deleting common and meaningless words such as "of", "and", "in"; text cleaning: removing special characters, punctuation marks, and irrelevant numbers; word segmentation: splitting the text into individual words or phrases for easy analysis;
[0074] Then extract keywords or phrases from the preprocessed business order text information as features:
[0075] Word frequency statistics: count the frequency of each word or phrase in the text; TF-IDF: use the term frequency-inverse document frequency method to evaluate the importance of a word for a single document in a text collection; N-gram model: extract combinations of N consecutive words to better capture word order information;
[0076] Then train a classification model using different machine learning algorithms. This model can classify order texts into different categories according to the extracted features. The classification model includes:
[0077] Naive Bayes: a simple probability-based classifier suitable for text classification; Support Vector Machine: a powerful classification algorithm suitable for high-dimensional data; Deep Learning: use convolutional neural networks or recurrent neural networks to capture complex patterns in text data;
[0078] Then apply the trained model to the preprocessed business order text information and classify the orders according to the keywords in the business order text information: by setting a threshold, when the prediction probability of the model exceeds this threshold, the order is classified into the corresponding category; when an order belongs to multiple categories, use the multi-label classification method;
[0079] Then evaluate the performance of the classification model and optimize it based on the feedback. The classification model is evaluated by obtaining the accuracy, recall, precision and F1 score of the classification model after classification. The accuracy is the proportion of orders correctly classified by the model, which is a model performance evaluation indicator. It means the proportion of orders correctly classified by the model to the total number of orders. The calculation formula is: , Z is the accuracy, q and s are the number of correctly classified orders and the total number of orders respectively;
[0080] The recall rate measures the ability of the model to correctly identify all orders. It measures the proportion of positive classes (related orders) correctly identified by the model to all actual positive classes. The calculation formula is: , where H is the recall rate, tp and fn are true positive examples and false negative examples respectively;
[0081] Precision measures the proportion of orders that are actually positive among those identified as positive by the model. The calculation formula is: , where J is the precision and fp is the false positive example;
[0082] The F1 score is an indicator that reconciles precision and recall. It is used to measure the overall performance of the model. The calculation formula is: , where F1 is the F1 score. The larger the calculated F1 score is, the better the performance of the corresponding classification model is.
[0083] Then normalize the obtained accuracy, recall, precision and F1 score and enter them into the following formula: To get the comprehensive evaluation value ZPZ, where are the preset weight coefficients for accuracy, recall, precision, and F1 score, and their values are 0.92, 1.05, 1.08, and 1.13, respectively;
[0084] The calculated comprehensive evaluation value ZPZ is then compared with the preset comprehensive evaluation threshold. When the comprehensive evaluation value is lower than the preset comprehensive evaluation threshold, it is judged that the classification model effect does not meet the standard, and the model is optimized based on the feedback; the feedback sources include:
[0085] User feedback: Collect feedback directly from business teams and end users who use the classification model to gather insights on model prediction accuracy, user experience, and actual business impact; Business indicators: Monitor business indicators, including customer satisfaction, order processing time, and inventory turnover, to evaluate the impact of the classification model on business processes; Error analysis: Analyze cases where the model predicts errors to understand the model's weaknesses in specific situations; Market changes: Pay attention to changes in market trends and customer needs, which can affect the effectiveness of the classification model;
[0086] The specific optimization items are as follows:
[0087] Adjust model parameters: adjust the model's hyperparameters, including learning rate, tree depth, and regularization strength, to improve model accuracy based on performance indicators and user feedback; Feature engineering: review and improve the feature extraction and selection process based on error analysis and user feedback, including adding new features, modifying feature conversion methods, or removing irrelevant features; Model selection and integration: when it is determined that a single model cannot meet performance requirements, use alternative classification models or adopt ensemble learning methods, including random forests, gradient boosting machines, or neural network integration; Dataset update and expansion: update and expand the training data set based on market changes and newly collected user feedback, including adding new samples, rebalancing category distribution, or introducing new data sources; Continuous monitoring and evaluation: establish a continuous monitoring and evaluation mechanism, regularly check the performance of the classification model, and make adjustments based on the latest data and feedback;
[0088] The determination module is used to determine the type of business order based on the user's analysis requirements;
[0089] After completing the classification of business orders, obtain the user's business order query requirements, and extract key features based on the user's query requirements, including product types, transaction time, and promotional activities, etc. At the same time, extract corresponding features from each business order that has been classified;
[0090] Then, the user demand features and each classified business order features are converted into numerical feature vectors through word frequency, word frequency-inverse document frequency (TF-IDF) or other text vectorization methods, and the user demand features and the classified business order features are marked as A and B respectively; then the similarity between the user demand feature vector and each classified business order feature vector is calculated using a similarity calculation method; wherein the similarity calculation method includes:
[0091] Cosine similarity: For vectors A and B, the similarity between them is determined by measuring the cosine value of the angle between the two vectors. The formula is: , A·B is the dot product of the vectors, ‖A‖ and ‖B‖ are the Euclidean norms or lengths of the vectors respectively;
[0092] Euclidean distance; for vectors A and B, the Euclidean distance is the straight-line distance between A and B, and the formula is: , when the Euclidean distance is smaller, the similarity is higher;
[0093] Jaccard similarity: For two sets A and B, the Jaccard similarity is the size of the intersection of the two sets divided by the size of the union. The formula is: , |A ∩ B| is the number of elements in the intersection of sets A and B, |A ∪ B| is the number of elements in the union of sets A and B;
[0094] Then conduct a comprehensive analysis of the above three similarity calculation methods, specifically:
[0095] Each similarity score is standardized to the range of [0, 1]. Then, different weights are assigned to each similarity method according to business needs or feature importance. The final comprehensive similarity score is calculated using weighted average or simple average. The weights of the three methods are set to w1, w2, and w3 respectively, and their scores are S1, S2, and S3 respectively. The comprehensive similarity score is calculated using the following formula: ;
[0096] The similarity between the characteristics of user needs and the characteristics of different classified business orders is calculated respectively to obtain different comprehensive similarity scores, and the different comprehensive similarity scores are sorted according to size. The largest comprehensive similarity score is selected, and the corresponding classified business order type is used as the type of user business order query demand.
[0097] The analysis module is used to analyze the business order trends that users need to query;
[0098] First, the data of the business order types identified in the big data are preprocessed, including data cleaning: removing errors and duplicate records, integration: format unification, normalization: numerical standardization and processing of missing values, so as to facilitate subsequent analysis; statistical tests, correlation analysis and feature selection algorithms are used to identify the most predictive features of the preprocessed data, so as to determine the features that have an important impact on trend prediction;
[0099] Then establish a time series analysis model to predict future trends, where the time series models include autoregression, moving average, autoregression moving average and autoregression integrated moving average models; use random forest, support vector machine, gradient boosting machine and neural network machine learning algorithms to predict trends, and train and verify the models. Specifically, use historical order data to train the selected model, and verify the model's prediction performance through cross-validation or retaining a portion of the data as a test set. Evaluation indicators include mean square error, root mean square error and mean absolute error;
[0100] Use the trained model to predict future order trends, and predict short-term (such as daily, weekly) or long-term (such as monthly, quarterly) trends based on business needs; then establish a correction model based on the prediction results of the model, and use the correction model to correct the prediction results. The specific steps are as follows:
[0101] First, identify and collect key features that affect order volume, including seasonal factors, promotional activities, economic indicators, competitor behavior, market trends, changes in customer behavior, and product updates or inventory levels;
[0102] Then use crawler technology, application programming interface calls (API) and third-party data services to collect relevant data on the above key features and conduct preliminary analysis to understand the relationship between these features and order volume. Then perform feature engineering on the collected data, including feature extraction, transformation and selection, use time series analysis methods to extract seasonal components, or use statistical methods to determine which features are most correlated with order volume;
[0103] Then, a correction model is established. The output of the basic prediction model is adjusted according to the changes in key features through the correction model. The correction model is then trained using the same type of business order data and key feature data to ensure that the training data covers all possible situations so that the model can learn how to adjust the prediction results according to the changes in key features.
[0104] Finally, the output of the basic forecasting model is used as input, combined with key feature data, and the revised model is used to adjust the forecast results. The revised forecast results should be closer to the actual observed order volume.
[0105] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A business order analysis and prediction system based on big data, characterized by: include: The classification module is used to classify business orders based on text analysis methods. The specific steps are as follows: First, obtain the business order text information, perform data preprocessing on the business order text information to remove noise and standardize the data format; Then extract keywords or phrases as features from the preprocessed business order text information: Word frequency statistics: count the frequency of each word or phrase appearing in the text; Term Frequency-Inverse Document Frequency: Use the term frequency-inverse document frequency method to evaluate the importance of a word to a document in a text collection; N-gram model: extracts combinations of N consecutive words to better capture word order information; Then, three machine learning algorithms, Naive Bayes, Support Vector Machine, and Deep Learning, were used to train the classification model; The trained model is then applied to the pre-processed business order text information, and the orders are classified according to the keywords in the business order text information: by setting a threshold, when the model's predicted probability exceeds this threshold, the order is classified into the corresponding category; when the order belongs to multiple categories, a multi-label classification method is used; Then evaluate the performance of the classification model and optimize it based on the feedback; The determination module is used to determine the type of business order for the user's analysis requirements. The specific process is as follows: After completing the classification of the business orders, the user's business order query requirements are obtained, and key features are extracted according to the user's query requirements, and corresponding features are extracted from each business order that has been classified; Then, the user demand features and each classified business order features are converted into numerical feature vectors through word frequency, word frequency-inverse document frequency or other text vectorization methods, and the user demand features and the classified business order features are marked as A and B respectively; Then three similarity calculation methods are used to calculate the similarity between the user demand feature vector and each classified business order feature vector; Then, the above three similarity calculation methods are comprehensively analyzed, specifically: Each similarity score is standardized to the range of [0, 1]. Then, different weights are assigned to each similarity method according to business needs or the importance of features. The final comprehensive similarity score is calculated using weighted average or simple average. The weights of the three similarity calculation methods are set to w1, w2, and w3, respectively, and the scores are S1, S2, and S3, respectively. The comprehensive similarity score is calculated using the following formula: ; The similarity between the characteristics of the user's needs and the characteristics of the different classified business orders is calculated to obtain different comprehensive similarity scores, and the obtained different comprehensive similarity scores are sorted according to size, and the largest comprehensive similarity score is selected, and the corresponding classified business order type is used as the type of the user's business order query demand; The analysis module is used to analyze the business order trends that users need to query.
2. The business order analysis and prediction system based on big data according to claim 1 is characterized in that: The specific process of the classification module evaluating the classification model is as follows: The classification model is evaluated by obtaining the accuracy, recall, precision and F1 score of the classification model after classification. The accuracy rate is a measure of the proportion of orders correctly classified by the model. It is a model performance evaluation indicator, indicating the proportion of orders correctly classified by the model to the total number of orders. The calculation formula is: , where Z is the accuracy, q and s are the number of correctly classified orders and the total number of orders respectively; The recall rate measures the ability of the model to correctly identify all orders. It measures the proportion of positive classes correctly identified by the model to all actual positive classes. The calculation formula is: , where H is the recall rate, tp and fn are true positive examples and false negative examples respectively; The precision rate measures the proportion of orders that are actually positive among the orders identified as positive by the model. The calculation formula is: , where J is the precision and fp is the false positive example; The F1 score is an indicator that reconciles precision and recall. It is used to measure the overall performance of the model. The calculation formula is: , where F1 is the F1 score. The larger the calculated F1 score is, the better the performance of the corresponding classification model is. Then normalize the obtained accuracy, recall, precision and F1 score and enter them into the following formula: To get the comprehensive evaluation value ZPZ, where They are the preset weight coefficients for accuracy, recall, precision and F1 score respectively; The calculated comprehensive evaluation value ZPZ is then compared with the preset comprehensive evaluation threshold. When the comprehensive evaluation value is lower than the preset comprehensive evaluation threshold, it is judged that the classification model effect does not meet the standard, and the model is optimized based on the feedback.
3. The business order analysis and prediction system based on big data according to claim 2 is characterized in that: The specific operation steps of the classification module to optimize the model according to the feedback are as follows: Feedback sources include: User feedback: Gather feedback from business teams and end users who use the classification model on model prediction accuracy, user experience, and actual business impact; Business Indicators: Monitor business indicators, including customer satisfaction, order processing time, and inventory turnover; Error analysis: Analyze cases where the model made incorrect predictions; Market changes: Pay attention to changes in market trends and customer needs; The specific optimization items are as follows: Adjust model parameters: Based on performance indicators and user feedback, adjust the model's hyperparameters, including learning rate, tree depth, and regularization strength; Feature Engineering: Improve the feature extraction and selection process based on error analysis and user feedback, including adding new features, modifying feature conversion methods, or removing irrelevant features; Model selection and integration: When it is determined that a single model cannot meet the performance requirements, use alternative classification models or adopt ensemble learning methods, including random forests, gradient boosting machines, or ensembles of neural networks; Dataset update and expansion: Update and expand the training dataset based on market changes and newly collected user feedback, including adding new samples, rebalancing category distribution, or introducing new data sources; Continuous Monitoring and Evaluation: Establish an ongoing monitoring and evaluation mechanism to regularly check the performance of the classification model and make adjustments based on the latest data and feedback.
4. The business order analysis and prediction system based on big data according to claim 1 is characterized in that: The specific operation steps of the determination module using three similarity calculation methods to calculate the similarity between the user demand feature vector and each classified business order feature vector are as follows: Similarity calculation methods include: Cosine similarity: For vectors A and B, the similarity between them is determined by measuring the cosine value of the angle between the two vectors. The formula is: , A·B is the dot product of the vectors, ‖A‖ and ‖B‖ are the Euclidean norms or lengths of the vectors respectively; Euclidean distance; for vectors A and B, the Euclidean distance is the straight-line distance between A and B, and the formula is: , When the Euclidean distance is smaller, the similarity is higher; Jaccard similarity: For two sets A and B, the Jaccard similarity is the size of the intersection of the two sets divided by the size of the union. The formula is: , |A ∩ B| is the number of elements in the intersection of sets A and B, and |A ∪ B| is the number of elements in the union of sets A and B.
5. The business order analysis and prediction system based on big data according to claim 1 is characterized in that: The specific operation steps of the analysis module for analyzing the business order trend that the user needs to query are as follows: First, the data of the business order types identified in the big data are preprocessed, including data cleaning, integration, normalization and processing of missing values; statistical tests, correlation analysis and feature selection algorithms are used to identify features with predictive value on the preprocessed data, thereby determining features that have an impact on trend prediction; Then establish a time series analysis model to predict future trends. The time series models include autoregression, moving average, autoregression moving average and autoregression integrated sliding average models, and use random forest, support vector machine, gradient boosting machine and neural network machine learning algorithms to predict trends. Then the model is trained and validated. Specifically, the selected model is trained using historical order data, and the prediction performance of the model is verified by cross-validation or by reserving a portion of the data as a test set. The evaluation indicators include mean square error, root mean square error, and mean absolute error. Use the trained model to predict future order trends and predict short-term or long-term trends based on business needs; Then, a correction model is established based on the prediction results of the model, and the prediction results are corrected through the correction model.
6. The business order analysis and prediction system based on big data according to claim 5 is characterized in that: The specific operation steps of the analysis module to correct the prediction results by correcting the model are as follows: First, identify and collect key features that affect order volume, including seasonal factors, promotional activities, economic indicators, competitor behavior, market trends, changes in customer behavior, and product updates or inventory levels; Then, we use crawler technology, application programming interface calls and third-party data services to collect relevant data on the above key features and conduct preliminary analysis to obtain the relationship between these features and order volume; Then perform feature engineering on the collected data, including feature extraction, transformation and selection, using time series analysis methods to extract seasonal components, or using statistical methods to determine which features are most correlated with order volume; Establish a correction model, adjust the output of the basic prediction model according to the changes in key features through the correction model, and then use the same type of business order data and key feature data to train the correction model; Finally, the output of the basic prediction model is used as input, combined with key feature data, and a correction model is used to adjust the prediction results.
Citation Information
Patent Citations
Order allocation method, device and apparatus and storage medium
CN113469615A
Order processing method and device, equipment and storage medium
CN113869576A