Smart shop data processing method and device based on multi-dimensional data analysis

By combining multidimensional data analysis and deep learning models with clustering rule mining and Markov decision processes, the problem of insufficient in-depth insights into consumer behavior in smart stores has been solved, enabling the dynamic generation and optimization of personalized marketing strategies and improving the accuracy of operational decisions and return on investment.

CN121921054APending Publication Date: 2026-04-24ZHUJI HAOXIN E-COMMERCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUJI HAOXIN E-COMMERCE CO LTD
Filing Date
2025-11-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing smart store data analysis technologies lack in-depth insights into consumer behavior and the mining of spatiotemporal patterns, making it difficult to achieve accurate predictions and personalized recommendations, resulting in insufficient support for operational decision-making.

Method used

Through multidimensional data analysis, a dynamic strategy generation model is constructed using a cascaded autoencoder-long short-term memory network model and a deep Q-network algorithm. This model combines clustering-association rule mining and Markov decision processes to generate differentiated marketing strategies, which are then optimized through online feedback.

Benefits of technology

It enables multi-dimensional and in-depth analysis of consumer behavior and product lifecycle, generating personalized and dynamically adjustable marketing strategies, thereby improving the accuracy of operational decisions and return on investment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921054A_ABST
    Figure CN121921054A_ABST
Patent Text Reader

Abstract

The invention provides a smart store data processing method and device based on multi-dimensional data analysis. Relates to the technical field of smart retail and big data processing, and comprises the following steps: S1, collecting multi-source heterogeneous original data through a data collection device in a shop, and performing cleaning, duplicate removal and unified format conversion on the multi-source heterogeneous original data to obtain preprocessed data to be fused; and S2, carrying out association mapping and time sequence synchronization on the to-be-fused data, and constructing a unified multi-dimensional data set including consumers, commodities and space-time dimensions. According to the smart store data processing method and device based on multi-dimensional data analysis, it is guaranteed that only minimum modification is needed when a system accesses new hardware or a business scene, the operation return on investment is maximized while the marketing precision is improved, and the method and device are suitable for smart store environments of various scales and business forms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart retail and big data processing technology, specifically to a smart store data processing method and apparatus based on multidimensional data analysis. Background Technology

[0002] Currently, with the widespread application of IoT, cloud computing, and AI technologies, smart stores can collect multi-source heterogeneous data such as store traffic, transactions, inventory, environment, and member profiles in real time by deploying various sensors, mobile payment terminals, membership management systems, and online e-commerce platforms. At the same time, some retail enterprises have introduced big data platforms to conduct statistical analysis on key indicators such as customer traffic and sales, and present the operational status in the form of reports or dashboards, providing support for store management and simple promotional decisions.

[0003] However, the aforementioned existing technologies still have significant shortcomings: the analysis relies heavily on traditional statistics and rule engines, lacks in-depth insights into consumer behavior and the mining of spatiotemporal patterns, makes it difficult to make accurate predictions and personalized recommendations, and makes it difficult for operators to obtain comprehensive decision support covering a global perspective of customer flow, products, personnel, environment, etc. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a smart store data processing method and apparatus based on multidimensional data analysis, which solves the problem of how to improve the store's data depth analysis capabilities and personalized marketing capabilities through dynamic strategy generation models based on multidimensional data.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a smart store data processing method based on multidimensional data analysis, comprising: S1. Collect multi-source heterogeneous raw data through data collection devices in the store, and perform cleaning, deduplication and unified format conversion on the multi-source heterogeneous raw data to obtain preprocessed data to be fused; S2. Perform correlation mapping and time-series synchronization on the data to be fused to construct a unified multidimensional dataset including consumers, products, and spatiotemporal dimensions; S3. Consumer behavior features, product association features, and spatiotemporal evolution features are extracted from the unified multidimensional dataset using a deep learning model, and in-depth analysis results covering consumer preferences and product lifecycle are obtained through clustering-association rule mining and temporal evolution analysis. S4. The deep analysis results, the unified multidimensional dataset, and the optional operational action sequence are modeled as a Markov decision process, and the deep learning model is trained based on the deep Q-network algorithm to output dynamic marketing strategies for differentiated customer groups. S5. Deploy the dynamic marketing strategy online at the store operation end, collect execution feedback signals, and fine-tune the deep learning model online based on the execution feedback signals to achieve closed-loop adaptive optimization of the personalized marketing strategy.

[0006] Preferably, the data acquisition device includes a customer flow camera, an environmental sensor, a cashier system, and a membership management system, and the multi-source heterogeneous raw data includes collected customer flow, transaction, inventory, environmental, and member profiles.

[0007] Preferably, the cleaning process performs outlier detection based on the sample mean and standard deviation, and removes samples that exceed three times the standard deviation.

[0008] Preferably, the unified format conversion process normalizes the numeric fields by standardizing all timestamps to a unified time zone format.

[0009] Preferably, the deep learning model is a hybrid network composed of a cascaded autoencoder and a long short-term memory network, and its model formula is as follows: in, For the first The original input vector of each sample, This is the output vector reconstructed by the autoencoder; This is the previous hidden state vector of the sample. This is the vector of the next hidden state; represents the weight hyperparameters for the hidden state reconstruction error.

[0010] Preferably, the deep Q-network iterates the Q-value using the following Bellman update formula: in, The learning rate controls the update step size. For at any time Execute action The instant rewards obtained This is a discount factor used to weigh current rewards against expected future rewards. For the value estimation of the current state-action pair, Value estimation for the next state-action pair.

[0011] Preferably, the execution feedback signals include click-through rate, conversion rate, and changes in average order value.

[0012] A smart store data processing device based on multidimensional data analysis includes: The data acquisition module is used to acquire multi-source heterogeneous raw data from sources such as passenger flow cameras, environmental sensors, POS systems, and membership management systems, and to clean, deduplicatize, and standardize the format of the multi-source heterogeneous raw data. The fusion module is used to perform association mapping and time-series alignment on the preprocessed data based on unified entity identifiers and timestamps, generating a unified multidimensional dataset containing consumer, product, and spatiotemporal dimensions; The analysis module is used to extract consumer behavior features, product association features, and spatiotemporal evolution features based on the unified multidimensional dataset by calling the cascaded autoencoder-long short-term memory network model, and to obtain in-depth analysis results through clustering and association rule mining. The decision-making module is used to construct a Markov decision process by combining the deep analysis results with the optional operational action sequence, and output dynamic marketing strategies for differentiated customer groups based on the deep Q-network algorithm. An optimization module is used to deploy the strategy online, collect feedback signals, and adjust the deep Q-network model online based on the feedback signals to achieve closed-loop adaptive optimization.

[0013] This invention provides a smart store data processing method and apparatus based on multidimensional data analysis. It has the following beneficial effects: This smart store data processing method and device based on multidimensional data analysis overcomes the limitations of existing technologies such as "data silos" and "statistical analysis" through real-time acquisition and efficient fusion of multi-source heterogeneous data. Utilizing a deep learning strategy that combines a cascaded autoencoder-long short-term memory network model with clustering and association rule mining, it can perform multi-dimensional in-depth analysis of consumer behavior, product associations, and spatiotemporal evolution, significantly improving the ability to understand customer preferences and product lifecycles. Furthermore, based on a Markov decision process constructed from a deep Q-network, it generates differentiated and dynamically adjustable personalized marketing strategies in real time. Through online feedback and a multi-armed slot machine mechanism for closed-loop optimization, it achieves an end-to-end adaptive closed loop from data insight to strategy execution.

[0014] The device of this invention adopts a modular architecture design with clear responsibilities for each module. It supports both second-level response at the in-store edge and collaborative training and updates in conjunction with the cloud. The pluggable data interface and unified entity identification mechanism ensure that the system requires only minimal modifications when connecting new hardware or business scenarios, maximizing operational return on investment while improving marketing accuracy. It is suitable for smart store environments of all sizes and business formats. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the process of realizing the invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1

[0018] like Figure 1 As shown, this embodiment of the invention provides a smart store data processing method based on multidimensional data analysis, including: S1. Collecting multi-source heterogeneous raw data through data acquisition devices in the store, and performing cleaning, deduplication, and unified format conversion on the multi-source heterogeneous raw data to obtain preprocessed data to be fused. The data acquisition devices include customer flow cameras, environmental sensors, a POS system, and a membership management system. The multi-source heterogeneous raw data includes collected customer flow, transaction, inventory, environment, and member profiles.

[0019] The specific implementation method is as follows: Multi-source heterogeneous raw data acquisition: Customer flow cameras: Four cameras are installed at the store entrance and main passageway to capture visitor profiles and entry / exit times in real time, generating 60,000 raw customer flow logs.

[0020] Environmental sensors: Two temperature and humidity sensors and two air quality sensors were installed in the store. Samples were taken every minute, and a total of 20,160 environmental data points were recorded.

[0021] The POS system: 6 POS machines recorded 12,345 transaction details.

[0022] Membership Management System: There are 5,200 online and offline members, generating 8,734 events such as check-in, points changes, and coupon redemption.

[0023] Outlier detection and removal: 1200 records with extremely low confidence or missing timestamps were removed from the passenger flow logs.

[0024] 960 consecutive duplicate records were removed from the environmental monitoring data.

[0025] A total of 445 "empty transactions" automatically generated in the POS system due to testing or refunds were removed.

[0026] After cleaning, 58,800 customer flow records, 19,200 environmental records, 11,900 transactions, and 8,734 member events were retained.

[0027] Data format standardization: Time unification: Convert all local times to UTC format.

[0028] Numerical mapping: Mapping continuous indicators such as camera confidence, temperature and humidity to the 0-1 range.

[0029] Category coding: Each member operation type is assigned an independent identifier to achieve distinguishable category characteristics.

[0030] Structured output: The logs from each device are parsed into a unified table format with consistent fields, which can be directly input into the subsequent fusion module.

[0031] Results: After completing the above cleaning, deduplication and format conversion, 98,634 preprocessed records were obtained. The structure was standardized and the fields were uniform, which laid the foundation for subsequent association mapping and in-depth analysis.

[0032] S2. Perform correlation mapping and time-series synchronization on the data to be merged to construct a unified multidimensional dataset that includes consumers, products, and spatiotemporal dimensions.

[0033] The cleaning process performs outlier detection based on the sample mean and standard deviation, removing samples exceeding three times the standard deviation. The unified format conversion process normalizes numerical fields by standardizing all timestamps to a unified timezone format.

[0034] The specific implementation method is as follows: Input data size: A total of 98,634 data entries were obtained from the S1 preprocessing process and are to be fused. Passenger flow log entries: 58,800.

[0035] 19,200 environmental records.

[0036] Transaction details: 11,900 transactions.

[0037] There were 8,734 member-related incidents.

[0038] Association mapping: Consumer-transaction mapping: Using the member IDs registered in the member management system, 11,310 valid member transactions out of 11,900 transactions were matched one-to-one with the profiles of 5,200 members, with a mapping coverage rate of 95%.

[0039] Transaction-Product Mapping: Based on the inventory management table, the 2,450 product numbers involved in 11,900 transactions are associated with the product attribute table to complete information such as product category, brand, and unit price.

[0040] Consumer-Customer Flow Mapping: Using a globally unique entity identifier (GUID), 1,250 in-store entry records with member IDs from 58,800 customer flow events are mapped to the corresponding consumer profiles.

[0041] Timing synchronization: Using a 5-minute sliding window, all timestamped events are aligned to UTC time: 11,900 transactions were mapped to the start time of the window, and a total of 2,376 time windows were mapped.

[0042] The 58,800 passenger flow records and 19,200 environmental data records are categorized into the same window, with an average of 9.4 passenger flow records and 2.5 environmental data records per window.

[0043] For each transaction, the customer flow and environmental data within the window are aggregated using either averaging or summing methods: Average passenger density: 9.4.

[0044] The average temperature is 24.6°C and the average humidity is 48.2%.

[0045] Constructing a unified multidimensional dataset: Ultimately, 11,900 unified records were generated, each containing: Consumer ID.

[0046] Transaction number and product details.

[0047] Window start time UTC.

[0048] Total number of passengers entering the window.

[0049] Average temperature and humidity inside the window.

[0050] Member profile characteristics.

[0051] Product attributes.

[0052] Store ID.

[0053] Results: Through the above-mentioned correlation mapping and time-series synchronization, not only was the precise alignment of 98,634 heterogeneous data points in the three dimensions of "consumer-product-spatiotemporal", but the originally scattered transaction, customer flow and environmental information was also integrated into 11,900 multi-dimensional analysis samples containing eight attributes, providing a high-quality unified input for subsequent deep feature mining.

[0054] S3. Consumer behavior features, product association features, and spatiotemporal evolution features are extracted from a unified multidimensional dataset using a deep learning model. In-depth analysis results covering consumer preferences and product lifecycle are obtained through clustering-association rule mining and temporal evolution analysis.

[0055] The deep learning model is a hybrid network consisting of a cascaded autoencoder and a long short-term memory network. Its model formula is: in, For the first The original input vector of each sample, This is the output vector reconstructed by the autoencoder; This is the previous hidden state vector of the sample. This is the vector of the next hidden state; represents the weight hyperparameters for the hidden state reconstruction error.

[0056] The specific implementation method is as follows: Model input and training: Input dimensions: Each record contains 8 fields.

[0057] Sample size: 11,900 samples in total, divided into a training set of 9,520 samples and a validation set of 2,380 samples.

[0058] Network structure: The autoencoder section consists of input → 4 → 2 → 4 → output, with hidden layer sizes of 4D and 2D respectively.

[0059] LSTM part: 2 layers stacked, with 16 hidden units in each layer.

[0060] Training process: Batch size 64, training 200 rounds.

[0061] The hidden state reconstruction error weight hyperparameter β is set to 0.5.

[0062] The reconstruction error after training converged to an average of 0.012 on the validation set.

[0063] Feature extraction: Consumer behavior characteristics: taken from the 16-dimensional hidden state vector at the last moment of the LSTM, used to represent the customer's activity level and preferences.

[0064] Product association features: taken from the intermediate bottleneck layer of the autoencoder—the potential association score of each product in each time period.

[0065] Spatiotemporal evolution characteristics: The hidden states of the same consumer in 5 consecutive sliding windows are spliced ​​together to obtain an evolution vector of 5×16=80 dimensions, which reflects the behavioral trend in a short period of time.

[0066] Clustering and Association Rule Mining: K-means clustering was performed on all consumer behavior characteristics, with K=5.

[0067] The sample proportions for each cluster were approximately 22%, 18%, 20%, 19%, and 21%, respectively.

[0068] Association rule mining is performed on the traded product combinations within each cluster, with a minimum support of 10% and a minimum confidence of 60%. Typical rules include: [Beverages → Snacks] Support rate: 12%, Confidence rate: 68%.

[0069] [Daily Necessities → Home Furnishings] Support rate: 11%, Confidence rate: 62%.

[0070] Time series evolution analysis: Based on spatiotemporal evolution vectors, similarity measurements were performed on the evolution sequences of customers within the same cluster. It was found that the purchase activity during the "14:00–16:00" window on weekdays was about 25% higher than on weekends.

[0071] For highly active clusters, it can also identify that the peak sales frequency of beer products after 18:00 on Fridays is as high as 35%.

[0072] In-depth analysis results include: five consumer profiles and their percentages, typical product pairing rules for each consumer group, and purchase popularity curves for different time periods and regions.

[0073] S4. The deep analysis results, unified multidimensional dataset, and optional operational action sequences are modeled as Markov decision processes, and a deep learning model is trained based on the deep Q-network algorithm to output dynamic marketing strategies for differentiated customer groups.

[0074] Deep Q-networks iterate the Q-value using the following Bellman update formula: in, The learning rate controls the update step size. For at any time Execute action The instant rewards obtained This is a discount factor used to weigh current rewards against expected future rewards. For the value estimation of the current state-action pair, Value estimation for the next state-action pair.

[0075] S5. Deploy dynamic marketing strategies online at the store operation level, collect execution feedback signals, and fine-tune the deep learning model online based on these signals to achieve closed-loop adaptive optimization of personalized marketing strategies. Execution feedback signals include changes in click-through rate, conversion rate, and average order value.

[0076] The specific implementation method is as follows: Application scenario: Dynamic marketing strategies during peak promotional periods.

[0077] State definition: The results of various in-depth analyses during the promotion period, along with factors such as consumption records, real-time inventory, and membership levels, are jointly encoded into the current state vector.

[0078] Action Space: Set up four types of actions: "Send limited-time discount push", "Add bonus points coupon", "Bundled recommendation", and "Live shopping guide", and specify the target audience for each action.

[0079] Model training: The deep Q-network is trained offline using historical "promotional period" data, prioritizing actions with high immediate returns while also considering future repurchase growth potential. By adjusting the learning rate and discount factor, the model can converge quickly in high-traffic environments.

[0080] Output strategy: Generate dynamic pricing and push timing plans for each differentiated user group. For example, focus on "live shopping guide + limited-time discount" for highly active users, and focus on "bonus points coupons" for potential churn users.

[0081] Deployment and execution: Publish the strategy to the store's app, WeChat official account, and in-store digital screens to achieve second-level push notifications.

[0082] Feedback collection: Real-time collection of metrics such as user click-through rate, order conversion rate, and changes in average order value.

[0083] Online fine-tuning: Within the multi-armed slot machine framework, the ROI of each action is evaluated in 10-minute cycles. Actions that significantly improve click-through rate and conversion rate are explored more frequently, while actions with poor feedback are quickly reduced in execution weight to ensure continuous optimization of return on investment during peak promotional periods.

[0084] A smart store data processing device based on multidimensional data analysis includes: The data acquisition module is used to acquire multi-source heterogeneous raw data from sources such as customer flow cameras, environmental sensors, POS systems, and membership management systems, and to clean, deduplicatize, and standardize the format of the multi-source heterogeneous raw data.

[0085] The fusion module is used to perform association mapping and time-series alignment on the preprocessed data based on unified entity identifiers and timestamps, generating a unified multidimensional dataset containing consumer, product, and spatiotemporal dimensions.

[0086] The analysis module is used to extract consumer behavior features, product association features, and spatiotemporal evolution features by calling the cascaded autoencoder-long short-term memory network model based on a unified multidimensional dataset, and to obtain in-depth analysis results through clustering and association rule mining.

[0087] The decision-making module is used to construct a Markov decision process by combining the results of in-depth analysis with the sequence of optional operational actions, and output dynamic marketing strategies for differentiated customer groups based on the deep Q-network algorithm.

[0088] The optimization module is used to deploy strategies online, collect feedback signals, and adjust the deep Q-network model online based on the feedback signals to achieve closed-loop adaptive optimization.

[0089] Example 2

[0090] Unlike Example 1, this example applies to an adaptive marketing strategy during a stable daily period.

[0091] State definition: The state vector is formed by combining information such as "people with high repurchase rates on weekday mornings" and "high-frequency small-amount purchase characteristics during lunch breaks" obtained from in-depth analysis with daily transaction and environmental data.

[0092] Action Space: We have selected three low-cost actions: "Limited-time free black tea coupons at lunchtime", "Member-exclusive discounts", and "Product recommendation cards".

[0093] Model training: Fine-tuning the deep Q-network on historical data during the plateau period, emphasizing cost sensitivity, and prioritizing low-cost actions while ensuring basic conversion.

[0094] Output strategy: Send "free black tea coupons" to customers who frequently visit the store during lunch breaks, send "exclusive discounts" to high-value members, and send "trial recommendation cards" to customers interested in new products.

[0095] Deployment and Implementation: The strategy is implemented through app push notifications and in-store self-service coupon machines, reaching the target audience at low cost.

[0096] Feedback collection: Summarize click-through rate, conversion rate, and average order value changes on a daily basis, with a focus on free coupon redemption rate and new member conversion rate.

[0097] Online fine-tuning: Evaluate the performance of the three actions daily, prioritize increasing the execution ratio of "free tea coupons" with high redemption rates but low costs, and appropriately reduce the execution ratio of "exclusive discounts" with high costs and average returns, so as to achieve controllable marketing costs and steady improvement in conversion rates during the stable period.

[0098] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A smart store data processing method based on multidimensional data analysis, characterized in that: include: S1. Collect multi-source heterogeneous raw data through data collection devices in the store, and perform cleaning, deduplication and unified format conversion on the multi-source heterogeneous raw data to obtain preprocessed data to be fused; S2. Perform correlation mapping and time-series synchronization on the data to be fused to construct a unified multidimensional dataset including consumers, products, and spatiotemporal dimensions; S3. Consumer behavior features, product association features, and spatiotemporal evolution features are extracted from the unified multidimensional dataset using a deep learning model, and in-depth analysis results covering consumer preferences and product lifecycle are obtained through clustering-association rule mining and temporal evolution analysis. S4. The deep analysis results, the unified multidimensional dataset, and the optional operational action sequence are modeled as a Markov decision process, and the deep learning model is trained based on the deep Q-network algorithm to output dynamic marketing strategies for differentiated customer groups. S5. Deploy the dynamic marketing strategy online at the store operation end, collect execution feedback signals, and fine-tune the deep learning model online based on the execution feedback signals.

2. The smart store data processing method based on multidimensional data analysis according to claim 1, characterized in that: The data acquisition equipment includes a customer flow camera, an environmental sensor, a POS system, and a membership management system. The multi-source heterogeneous raw data includes collected customer flow, transaction, inventory, environmental, and member profiles.

3. The smart store data processing method based on multidimensional data analysis according to claim 1, characterized in that: The cleaning process performs outlier detection based on the sample mean and standard deviation, removing samples that exceed three times the standard deviation.

4. The smart store data processing method based on multidimensional data analysis according to claim 1, characterized in that: The unified format conversion process normalizes numerical fields by standardizing all timestamps to a unified time zone format.

5. The smart store data processing method based on multidimensional data analysis according to claim 1, characterized in that: The deep learning model is a hybrid network consisting of a cascaded autoencoder and a long short-term memory network, and its model formula is as follows: in, For the first The original input vector of each sample, This is the output vector reconstructed by the autoencoder; This is the previous hidden state vector of the sample. This is the vector of the next hidden state; represents the weight hyperparameters for the hidden state reconstruction error.

6. The smart store data processing method based on multidimensional data analysis according to claim 1, characterized in that: The deep Q-network iterates the Q-value using the following Bellman update formula: in, For learning rate, For at any time Execute action The instant rewards received As a discount factor, For the value estimation of the current state-action pair, Value estimation for the next state-action pair.

7. The smart store data processing method based on multidimensional data analysis according to claim 1, characterized in that: The execution feedback signals include click-through rate, conversion rate, and changes in average order value.

8. A smart store data processing device based on multidimensional data analysis, characterized in that: include: The data acquisition module is used to acquire multi-source heterogeneous raw data from sources such as passenger flow cameras, environmental sensors, POS systems, and membership management systems, and to clean, deduplicatize, and standardize the format of the multi-source heterogeneous raw data. The fusion module is used to perform association mapping and time-series alignment on the preprocessed data based on unified entity identifiers and timestamps, generating a unified multidimensional dataset containing consumer, product, and spatiotemporal dimensions; The analysis module is used to extract consumer behavior features, product association features, and spatiotemporal evolution features based on the unified multidimensional dataset by calling the cascaded autoencoder-long short-term memory network model, and to obtain in-depth analysis results through clustering and association rule mining. The decision-making module is used to construct a Markov decision process by combining the deep analysis results with the optional operational action sequence, and output dynamic marketing strategies for differentiated customer groups based on the deep Q-network algorithm. An optimization module is used to deploy the strategy online, collect feedback signals, and adjust the deep Q-network model online based on the feedback signals.