An e-commerce information storage method and system based on big data analysis

By improving the fruit fly optimization algorithm and locality-sensitive hashing technology, a hierarchical storage structure and multi-level index tables were constructed, solving the problems of low query efficiency and rudimentary management of hot and cold data in e-commerce data storage, and achieving efficient data distribution management and query optimization.

CN120277068BActive Publication Date: 2025-11-11BEIJING ZHISUANDUODUO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510333523.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-11-11
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Existing e-commerce data storage solutions struggle to meet high-performance requirements in massive data environments, suffer from low query efficiency, rudimentary management of hot and cold data, and fail to intelligently optimize based on data characteristics and access patterns. Consequently, storage structures lack specificity and fail to achieve efficient data distribution management.

Method used

An improved fruit fly optimization algorithm is used to perform global optimization and local tuning of e-commerce product data. By combining the access priority and historical access hotspot information of e-commerce product data, a hierarchical storage structure is constructed, and locality-sensitive hashing technology is introduced to generate multi-level index tables and dynamically adjust the data storage location.

Benefits of technology

It improves data access performance, reduces the load imbalance of storage nodes, optimizes query efficiency, and ensures the efficient operation of the system in a large-scale data environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277068B_ABST
    Figure CN120277068B_ABST
Patent Text Reader

Abstract

This invention discloses an e-commerce information storage method and system based on big data analysis. The method comprises: S1. Formatting e-commerce product data collected from an e-commerce platform to form a standardized e-commerce product dataset; S2. Dividing the e-commerce product data into different storage areas according to the access frequency, data type, and storage resource allocation requirements; S3. Forming an optimized e-commerce product data storage scheme; S4. Establishing a unified multi-level local sensitive hash index table for e-commerce product data in each storage area; S5. Combining the optimized e-commerce product data storage scheme with the e-commerce product data index system to form a unified e-commerce information storage management system. This invention enables dynamic adjustment of e-commerce product data storage between hot and cold storage areas, avoiding the storage of hot data in inefficient storage areas and improving access performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information storage technology, and in particular to an e-commerce information storage method and system based on big data analysis. Background Technology

[0002] With the rapid development of the e-commerce industry, the product data on e-commerce platforms has exploded. Multi-dimensional information such as product types, transaction records, and user interaction data is constantly accumulating. The storage and management of this large-scale data has become one of the key issues in the operation of e-commerce platforms. The storage of e-commerce information not only requires efficient access capabilities to ensure users' rapid query and response, but also needs to optimize the data storage structure to reduce storage costs and improve the scalability of data management. In the current technology, the storage of e-commerce product data mainly adopts traditional relational databases or distributed storage solutions, but these solutions have obvious limitations.

[0003] Currently, traditional e-commerce data storage methods typically rely on relational databases or NoSQL databases. While relational databases have mature management mechanisms for structured data storage, their query performance gradually declines as the volume of e-commerce data grows. Especially in massive data environments, the overhead of index maintenance is too high, affecting the system's response speed. In addition, relational databases have poor scalability and cannot support the needs of e-commerce platforms for large-scale concurrent access.

[0004] To alleviate the storage and query performance bottlenecks of relational databases, many e-commerce platforms have introduced NoSQL databases and distributed storage systems in recent years. However, since e-commerce data usually contains a large amount of historical transaction records and high-dimensional data such as user interaction information, NoSQL databases still have shortcomings in handling complex queries and transaction consistency. In addition, most existing NoSQL databases lack efficient index optimization mechanisms, which makes it difficult to meet the high-performance requirements of query efficiency in large-scale data environments.

[0005] Furthermore, with the diversified data storage needs of e-commerce platforms, the tiered management of hot and cold data access has become one of the key technologies for optimizing storage. Some current data storage optimization solutions are mainly based on the access frequency or time characteristics of data, dividing data into cold and hot data and using different storage media for tiered management. However, the simple hot and cold data division method has the following problems: it cannot make full use of the type characteristics of product data and the storage resource allocation requirements, resulting in a lack of targeted optimization of the storage structure; the data migration strategy is relatively crude and fails to combine dynamic access mode for fine management, which affects data access efficiency.

[0006] In addition, existing e-commerce data storage optimization methods still have certain shortcomings in terms of global optimization and local tuning. For example, some systems use heuristic algorithms or traditional optimization methods to adjust the data storage structure, but because they fail to fully combine the access priority, storage resource requirements and load balancing characteristics of e-commerce data, it is often difficult to optimize storage performance while ensuring the efficiency of data access.

[0007] In summary, existing technologies for e-commerce product data storage mainly suffer from the following problems: First, traditional database solutions struggle to meet the high-performance demands of massive data storage and retrieval; second, existing NoSQL storage solutions lack efficient index optimization mechanisms, resulting in limited query efficiency; third, existing cold and hot data management methods are relatively crude, failing to combine data characteristics and access patterns for intelligent optimization; and fourth, existing storage optimization algorithms are insufficient in global and local tuning, making it difficult to achieve efficient data distribution management. Therefore, there is an urgent need for an e-commerce information storage method based on big data analytics to improve data storage flexibility, query efficiency, and optimization effectiveness. Summary of the Invention

[0008] One objective of this invention is to propose an e-commerce information storage method and system based on big data analysis. This invention enables e-commerce product data to be dynamically adjusted between hot and cold storage areas, avoiding hot data storage in inefficient storage areas and improving access performance.

[0009] According to an embodiment of the present invention, an e-commerce information storage method and system based on big data analysis includes the following steps:

[0010] S1. Format the e-commerce product data collected from the e-commerce platform to form a standardized e-commerce product dataset;

[0011] S2. Construct a hierarchical storage structure based on the standardized e-commerce product dataset, divide the e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of the e-commerce product data, and preset the access priority of each storage area;

[0012] S3. The improved fruit fly optimization algorithm is used to perform global search and local optimization on the hierarchical storage structure to optimize the distribution of e-commerce commodity data in each storage area and the resource allocation of storage nodes, thus forming an optimized e-commerce commodity data storage scheme.

[0013] S4. The standardized e-commerce product dataset is constructed using locality-sensitive hashing (LSH) technology to create an e-commerce product data index. A low-dimensional hash index is generated based on the high-dimensional feature information of the e-commerce product data, and a unified multi-level LSH index table for e-commerce product data is established in each storage area.

[0014] S5. The optimized e-commerce commodity data storage scheme is combined with a multi-level local sensitive hash index table to form a unified e-commerce information storage management system.

[0015] Optionally, S1 includes the following steps:

[0016] S11. Collect e-commerce product data from the e-commerce platform, wherein the e-commerce product data includes basic product information, classification information, historical transaction records and user interaction information, and construct an e-commerce product dataset;

[0017] S12. Format the e-commerce product dataset so that all data items in the e-commerce product dataset meet a unified data storage format standard;

[0018] S13. Perform a consistency check on the formatted e-commerce product dataset to ensure that all data items within the dataset meet the requirements of integrity, correctness, and uniqueness. Define the formatted standardized e-commerce product dataset as follows:

[0019] ;

[0020] in, For the standardized e-commerce product dataset, For the formatted first Each product data record , , , These respectively represent standardized basic product information, product category information, historical transaction records, and user interaction information. This represents the total amount of product data.

[0021] Optionally, S2 includes the following steps:

[0022] S21. Based on the standardized e-commerce product dataset To analyze the access frequency, data type, and storage resource requirements of e-commerce product data, an e-commerce product data feature matrix is ​​established.

[0023] ;

[0024] in, For e-commerce product data feature matrix, Indicates the first A set of features for e-commerce product data records. For the first The frequency of access to e-commerce product data For the first Data types of e-commerce product data For the first Storage resource requirements for e-commerce product data;

[0025] S22. Based on the e-commerce product data feature matrix Determine the storage area partitioning rules for e-commerce product data, and map the e-commerce product data to different storage areas:

[0026] ;

[0027] in, This represents a collection of data storage areas for e-commerce product information. For the first One storage area, Indicates the total number of storage areas. , , They represent the first The access frequency range, data type range, and storage resource allocation rules corresponding to each storage region;

[0028] S23. For each storage region Preset access priorities, defining the access priorities for each storage area as follows:

[0029] ;

[0030] in, Indicates the first Access priority of each storage region , , The first Normalized average values ​​of e-commerce product data access frequency, data type, and storage resource requirements within each storage region. , , These are the weighting coefficients for access frequency, data type, and storage resource requirements, respectively.

[0031] S24. Based on storage region access priority For the set of storage regions Sort the data to complete the hierarchical storage structure of e-commerce product data.

[0032] Optionally, S3 includes the following steps:

[0033] S31. Based on the hierarchical storage structure of e-commerce product data, and with the goals of reducing e-commerce platform response latency, reducing data redundancy, and improving load balancing of storage nodes, establish an objective function for optimizing e-commerce product data storage:

[0034] ;

[0035] in, Optimize the objective function for e-commerce product data storage. For the first storage areas The average access latency for e-commerce product data in China This represents the maximum latency for accessing e-commerce product data in each storage region. For the first The proportion of redundant data in e-commerce product data in each storage area This represents the maximum redundancy ratio of e-commerce product data in each storage region. For the first The degree of load imbalance of storage nodes in each storage region This represents the maximum value of the load imbalance between storage nodes in each storage region. , , These are the optimization weighting coefficients for access latency, data redundancy, and load balancing in e-commerce platforms, respectively.

[0036] S32. Initialize the population of the improved fruit fly optimization algorithm suitable for optimizing the data storage structure of e-commerce products, and define the initial fruit fly individual position as follows:

[0037] ;

[0038] in, This is the initial set of individual locations in the fruit fly population. For the first One fruit fly individual, For population size, Indicates the first Distribution scheme of individual e-commerce product data across various storage areas Indicates the first Initial configuration scheme for storage node resources corresponding to each individual;

[0039] S33. Combining an adaptive sniffing guidance strategy for prioritizing access to e-commerce product data storage, with the objective function of e-commerce product data storage optimization. To dynamically update the location of individual fruit flies:

[0040] ;

[0041] in, For the next iteration Individual fruit fly location update scheme The global optimization sensitivity coefficient, For access priority sensitive factors, For the first Storage area access priority, This represents the optimal storage location scheme in the current iteration. This is the random disturbance adjustment coefficient. To combine storage areas Lévy flight random step size for access priority adjustment For the first The objective function value for optimizing e-commerce product data storage for each individual entity. Optimize the objective function value for storing the optimal e-commerce product data in the current population;

[0042] S34. Based on historical access hotspot information from the e-commerce platform, construct a memory-enhanced local clustering optimization strategy to further update the individual fruit fly positions:

[0043] ;

[0044] in, A memory-enhanced e-commerce product data storage location scheme. Sensitivity coefficient for historical hot topics Indicates the first Each individual corresponds to a storage area Historical access frequency, For the first Historical average access frequency of each storage region For storage area The local optimal storage location is determined by the access frequency of historical hot e-commerce product data.

[0045] S35. Repeat steps S33 and S34 above, and dynamically adjust the global optimization sensitivity coefficient according to the real-time adaptability changes of the e-commerce commodity data storage structure. Random disturbance adjustment coefficient Sensitivity coefficient to historical hot topics The scheme for the location of individual fruit fly populations is continuously iterated and updated until the objective function for optimizing e-commerce commodity data storage is optimized. Convergence, output the final optimized e-commerce product data storage solution:

[0046] ;

[0047] in, This represents the final optimized scheme for the location of individual fruit fly individuals in the population. To optimize the storage distribution of e-commerce product data across various storage areas, This is the optimized storage node resource configuration scheme.

[0048] Optionally, S4 includes the following steps:

[0049] S41. Based on standardized e-commerce product datasets Extract high-dimensional feature information from e-commerce product data and construct a high-dimensional feature matrix suitable for indexing e-commerce product data:

[0050] ;

[0051] in, This represents a high-dimensional feature matrix of e-commerce product data. For the first A high-dimensional feature vector of e-commerce product data Indicates the first The first e-commerce product data 3D eigenvalues The total dimension of the high-dimensional features. The total amount of e-commerce product data;

[0052] S42. Based on the historical access frequency characteristics of e-commerce platform product data, an adaptive hash weight adjustment strategy for e-commerce product data access popularity is introduced. A method for generating locally sensitive hash index vectors for e-commerce product data that adapts to access popularity is constructed, defining the... The hash index vector of each e-commerce product data is:

[0053] ;

[0054] in, For the first A low-dimensional hash index vector of e-commerce product data. For the first The first e-commerce product data There are several hash index values, where L is the dimension of the Local Sensitive Hash Index Vector for e-commerce product data. For the first The initial weight vector of random projections of hash functions. The first [item] obtained through adaptive learning based on historical access hotspot features of e-commerce product data. Adjust the weight vector for each hot spot. For the first The frequency of access to e-commerce product data This is the adjustment coefficient for the hash weight adjustment based on the access popularity of the e-commerce platform. For the first The offset value of each hash function;

[0055] S43. Based on the collection of e-commerce commodity data storage areas Calculate each storage region The characteristic mean and covariance matrix of the data:

[0056] ;

[0057] ;

[0058] in, Indicates storage area The average feature vector of domestic e-commerce product data, For storage area The covariance matrix of domestic e-commerce product data For storage area The quantity of e-commerce product data within the country, where T is the transpose;

[0059] Furthermore, it optimizes the accuracy of approximate retrieval of e-commerce product data by defining an adaptive hash distance that combines differences in e-commerce product data types and storage area access priorities:

[0060] ;

[0061] in, For the first The and the first The hash distance of e-commerce product data , E-commerce product data and The corresponding number Each hash index value To combine the differences in e-commerce product data types with the first storage areas The access priority is adaptively determined by the first Weight adjustment factor for each hash index bit:

[0062] ;

[0063] in, Adjusting the weights for hash distance calculation based on storage region characteristics. and Storage areas Inner The mean and variance of the dimensional features;

[0064] S44. Based on the hash index vector set and adaptive hash distance constructed in steps S42 and S43, construct a multi-level locality-sensitive hash index table suitable for e-commerce product data location and approximate retrieval:

[0065] ;

[0066] in, For the first A multi-level locality-sensitive hash index table for e-commerce product data in each storage area. For the first A hash index vector of e-commerce product data. For e-commerce product data in the storage area The storage address in This refers to the hash index hierarchy that is divided within the storage area based on the access frequency and data type of e-commerce product data.

[0067] Optionally, S5 includes the following steps:

[0068] S51. E-commerce Commodity Data Storage Solution With multi-level locality sensitive hash index table Establish a global mapping relationship for the e-commerce information storage management system, and define a global data mapping relationship set:

[0069] ;

[0070] in, This refers to the set of global data mapping relationships within the e-commerce information storage and management system. This is the hash index vector corresponding to the e-commerce product data. This is the physical storage address for the e-commerce product data. This refers to the hierarchical information of the e-commerce product data in the multi-level locality-sensitive hash index table;

[0071] S52. Based on the dynamic access frequency of e-commerce product data, dynamically adjust the data storage location of e-commerce product data and define storage migration optimization rules:

[0072] ;

[0073] in, The adjusted location for storing e-commerce product data. To store migration weight parameters, For the first Real-time access priority for e-commerce product data. For storage area The historical average access priority;

[0074] S53. Adjust the hash index table according to changes in access patterns of e-commerce product data. Define the dynamic update rules for the hash index:

[0075] ;

[0076] in, For the first The updated hash index table at each time point. This indicates new index entries added due to new data or data storage migration. These are index deletions caused by data migration or deletion;

[0077] S54. Construct an access path optimization model for e-commerce product data by integrating the global storage management system for e-commerce product data. The e-commerce product data access path optimization model considers data storage location. Hash index table after index update Based on the combined effects, we define the objective function for optimizing e-commerce product data access:

[0078] ;

[0079] in, Optimize the objective function for accessing e-commerce product data. For the first Query response time for e-commerce product data. For the first The computational complexity of querying e-commerce product data. The extent to which index updates affect the query path:

[0080] ;

[0081] in, Adjust the weighting coefficients of the index's impact on the access optimization objective function;

[0082] The impact of index updates on query paths This reflects the impact of index updates on the data access path. The larger the value, the more drastic the index update, which will cause changes in the data query path.

[0083] S55. Based on the optimized data storage scheme Optimization objective function for e-commerce product data access Construct the final e-commerce information storage management system and perform iterative optimization of the storage structure:

[0084] ;

[0085] in, For the first The data of the first e-commerce product in the first The final storage location after the next iteration Adjust the step size factor for global storage. For the first The data of the first e-commerce product in the first The access optimization target value at the next iteration. and The first The minimum and maximum values ​​of the objective function for optimizing e-commerce product data access in each iteration. For the latest storage location, This is the final storage location from the previous round of optimization. Adjusting the impact factor of index updates on storage scheme updates, and considering the degree of impact of index updates on query paths. If the value exceeds the preset value, the storage location will be adjusted to adapt to the index changes and optimize retrieval efficiency.

[0086] An e-commerce information storage system based on big data analytics, used to execute an e-commerce information storage method based on big data analytics, includes the following modules:

[0087] The data preprocessing module is used to receive e-commerce product data from the e-commerce platform, and to standardize the data to generate a standardized e-commerce product dataset.

[0088] The hierarchical storage management module is used to build a hierarchical storage structure based on standardized e-commerce product datasets. According to the access frequency, data type and storage resource allocation requirements of e-commerce product data, it divides e-commerce product data into different storage areas and sets the access priority of each storage area.

[0089] An improved fruit fly optimization storage module is used to optimize the distribution of e-commerce product data in various storage areas and the resource configuration of storage nodes based on the improved fruit fly optimization algorithm, and to establish an optimized e-commerce product data storage scheme.

[0090] The index building and optimization module is used to build an e-commerce product data index on a standardized e-commerce product dataset using locality-sensitive hashing (LSH) technology. It generates a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data and establishes a multi-level LSH index table for the e-commerce product data.

[0091] The information storage management module is used to combine the optimized e-commerce commodity data storage scheme with the multi-level local sensitive hash index table of e-commerce commodity data to build a unified e-commerce information storage management system.

[0092] The beneficial effects of this invention are:

[0093] (1) The present invention uses an improved fruit fly optimization algorithm to perform global optimization and local tuning of the storage distribution of e-commerce commodity data. By combining the access priority of e-commerce commodity data, adaptive olfactory guidance strategy and historical access hotspot information, the optimal distribution of e-commerce commodity data in the storage area is achieved. Furthermore, access frequency sensitive factor and historical hotspot enhancement strategy are introduced, so that the fruit fly optimization algorithm can dynamically adapt to the access patterns of different e-commerce commodity data during the data distribution adjustment process, thereby reducing the load imbalance of storage nodes and improving the concurrent response capability of data access.

[0094] (2) This invention introduces locality-sensitive hashing technology and combines it with an adaptive weight adjustment strategy based on the access popularity of e-commerce product data to construct a multi-level hash index table. By adaptively adjusting the weight of the hash index, hot data can be preferentially mapped to high-priority storage areas, thereby improving query efficiency.

[0095] (3) The present invention constructs an adaptive hierarchical storage structure based on a standardized e-commerce product dataset, and adjusts the data storage location according to the real-time access mode of e-commerce data through a dynamic storage migration optimization strategy, thereby achieving a reasonable allocation of storage resources. Specifically, the storage migration optimization rule based on access priority enables e-commerce product data to be dynamically adjusted between cold and hot data storage areas, avoiding hot data storage in inefficient storage areas and improving access performance. Attached Figure Description

[0096] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0097] Figure 1 This is a flowchart of an e-commerce information storage method and system based on big data analysis proposed in this invention. Detailed Implementation

[0098] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0099] refer to Figure 1A method and system for storing e-commerce information based on big data analysis, comprising the following steps:

[0100] S1. Format the e-commerce product data collected from the e-commerce platform to form a standardized e-commerce product dataset;

[0101] S2. Construct a hierarchical storage structure based on standardized e-commerce product datasets. Divide e-commerce product data into different storage areas according to the access frequency, data type, and storage resource allocation requirements of the e-commerce product data, and preset the access priority of each storage area.

[0102] S3. Using the improved fruit fly optimization algorithm, a global search and local optimization of the hierarchical storage structure are performed to optimize the distribution of e-commerce commodity data in each storage area and the resource allocation of storage nodes, thus forming an optimized e-commerce commodity data storage scheme.

[0103] S4. Use locality-sensitive hashing (LSH) technology to construct an e-commerce product data index for the standardized e-commerce product dataset. Generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data, and establish a unified multi-level LSH index table for e-commerce product data in each storage area.

[0104] S5. Combine the optimized e-commerce product data storage scheme with a multi-level local sensitive hash index table to form a unified e-commerce information storage management system.

[0105] In this embodiment, S1 includes the following steps:

[0106] S11. Collect e-commerce product data from the e-commerce platform. The e-commerce product data includes basic product information, classification information, historical transaction records, and user interaction information, and construct an e-commerce product dataset.

[0107] S12. Format the e-commerce product dataset so that all data items in the e-commerce product dataset meet the unified data storage format standard;

[0108] S13. Perform a consistency check on the formatted e-commerce product dataset to ensure that all data items within the dataset meet the requirements of integrity, correctness, and uniqueness. Define the formatted standardized e-commerce product dataset as follows:

[0109] ;

[0110] in, For the standardized e-commerce product dataset, For the formatted first Each product data record , , , These respectively represent standardized basic product information, product category information, historical transaction records, and user interaction information. This represents the total amount of product data.

[0111] This implementation method standardizes e-commerce product data to ensure data format consistency, improve the reliability of data storage and retrieval, and adopts a structured data organization method to provide a unified storage standard for product information, transaction records, and user interaction data, reducing data redundancy and improving the accuracy of data queries. Furthermore, through formatting and consistency checks, it effectively enhances data integrity and correctness, reduces storage and calculation errors caused by data format mismatches, and provides a solid foundation for subsequent data storage optimization and index construction.

[0112] In this embodiment, S2 includes the following steps:

[0113] S21. Based on the standardized e-commerce product dataset To analyze the access frequency, data type, and storage resource requirements of e-commerce product data, an e-commerce product data feature matrix is ​​established.

[0114] ;

[0115] in, For e-commerce product data feature matrix, Indicates the first A set of features for e-commerce product data records. For the first The frequency of access to e-commerce product data For the first Data types of e-commerce product data For the first Storage resource requirements for e-commerce product data;

[0116] S22. Based on the e-commerce product data feature matrix Determine the storage area partitioning rules for e-commerce product data, and map the e-commerce product data to different storage areas:

[0117] ;

[0118] in, This represents a collection of data storage areas for e-commerce product information. For the first One storage area, Indicates the total number of storage areas. , , They represent the first The access frequency range, data type range, and storage resource allocation rules corresponding to each storage region;

[0119] S23. For each storage region Preset access priorities, defining the access priorities for each storage area as follows:

[0120] ;

[0121] in, Indicates the first Access priority of each storage region , , The first Normalized average values ​​of e-commerce product data access frequency, data type, and storage resource requirements within each storage region. , , These are the weighting coefficients for access frequency, data type, and storage resource requirements, respectively.

[0122] S24. Based on storage region access priority For storage area set Sort the data to complete the hierarchical storage structure of e-commerce product data.

[0123] This implementation method constructs a hierarchical storage structure based on the access frequency, data type, and storage resource allocation requirements of e-commerce product data, thereby improving the flexibility of data management and storage efficiency. By pre-setting access priorities for storage areas, frequently accessed data is stored in low-latency storage areas, optimizing data retrieval speed. Simultaneously, the hierarchical data storage approach reduces storage resource waste, improves system load balancing capabilities, and ensures that the e-commerce platform maintains efficient and stable operation even in large-scale data storage environments.

[0124] In this embodiment, S3 includes the following steps:

[0125] S31. Based on the hierarchical storage structure of e-commerce product data, and with the goals of reducing e-commerce platform response latency, reducing data redundancy, and improving load balancing of storage nodes, establish an objective function for optimizing e-commerce product data storage:

[0126] ;

[0127] in, Optimize the objective function for e-commerce product data storage. For the first storage areas The average access latency for e-commerce product data in China This represents the maximum latency for accessing e-commerce product data in each storage region. For the first The proportion of redundant data in e-commerce product data in each storage area This represents the maximum redundancy ratio of e-commerce product data in each storage region. For the first The degree of load imbalance of storage nodes in each storage region This represents the maximum value of the load imbalance between storage nodes in each storage region. , , These are the optimization weighting coefficients for access latency, data redundancy, and load balancing in e-commerce platforms, respectively.

[0128] S32. Initialize the population of the improved fruit fly optimization algorithm suitable for optimizing the data storage structure of e-commerce products, and define the initial fruit fly individual position as follows:

[0129] ;

[0130] in, This is the initial set of individual locations in the fruit fly population. For the first One fruit fly individual, For population size, Indicates the first Distribution scheme of individual e-commerce product data across various storage areas Indicates the first Initial configuration scheme for storage node resources corresponding to each individual;

[0131] S33. Combining an adaptive sniffing guidance strategy for prioritizing access to e-commerce product data storage, with the objective function of e-commerce product data storage optimization. To dynamically update the location of individual fruit flies:

[0132] ;

[0133] in, For the next iteration Individual fruit fly location update scheme The global optimization sensitivity coefficient, For access priority sensitive factors, For the first Storage area access priority, This represents the optimal storage location scheme in the current iteration. This is the random disturbance adjustment coefficient. To combine storage areas Lévy flight random step size for access priority adjustment For the first The objective function value for optimizing e-commerce product data storage for each individual entity. Optimize the objective function value for storing the optimal e-commerce product data in the current population;

[0134] S34. Based on historical access hotspot information from the e-commerce platform, construct a memory-enhanced local clustering optimization strategy to further update the individual fruit fly positions:

[0135] ;

[0136] in, A memory-enhanced e-commerce product data storage location scheme. Sensitivity coefficient for historical hot topics Indicates the first Each individual corresponds to a storage area Historical access frequency, For the first Historical average access frequency of each storage region For storage area The local optimal storage location is determined by the access frequency of historical hot e-commerce product data.

[0137] S35. Repeat steps S33 and S34 above, and dynamically adjust the global optimization sensitivity coefficient according to the real-time adaptability changes of the e-commerce commodity data storage structure. Random disturbance adjustment coefficient Sensitivity coefficient to historical hot topics The scheme for the location of individual fruit fly populations is continuously iterated and updated until the objective function for optimizing e-commerce commodity data storage is optimized. Convergence, output the final optimized e-commerce product data storage solution:

[0138] ;

[0139] in, This represents the final optimized scheme for the location of individual fruit fly individuals in the population. To optimize the storage distribution of e-commerce product data across various storage areas, This is the optimized storage node resource configuration scheme.

[0140] This implementation utilizes an improved fruit fly optimization algorithm to optimize e-commerce product data storage, achieving dynamic adaptive adjustments in storage node resource configuration and data distribution. Through an optimization strategy combining global search and local tuning, it ensures the data storage structure conforms to the optimal storage layout, improving query response speed. An adaptive access priority-based guidance strategy is introduced, enabling data storage optimization to effectively adapt to the characteristics of e-commerce data access, improving system storage performance, data access efficiency, and storage resource utilization, while reducing data access latency.

[0141] In this embodiment, S4 includes the following steps:

[0142] S41. Based on standardized e-commerce product datasets Extract high-dimensional feature information from e-commerce product data and construct a high-dimensional feature matrix suitable for indexing e-commerce product data:

[0143] ;

[0144] in, This represents a high-dimensional feature matrix of e-commerce product data. For the first A high-dimensional feature vector of e-commerce product data Indicates the first The first e-commerce product data 3D eigenvalues The total dimension of the high-dimensional features. The total amount of e-commerce product data;

[0145] S42. Based on the historical access frequency characteristics of e-commerce platform product data, an adaptive hash weight adjustment strategy for e-commerce product data access popularity is introduced. A method for generating locally sensitive hash index vectors for e-commerce product data that adapts to access popularity is constructed, defining the... The hash index vector of each e-commerce product data is:

[0146] ;

[0147] in, For the first A low-dimensional hash index vector of e-commerce product data. For the first The first e-commerce product data There are several hash index values, where L is the dimension of the Local Sensitive Hash Index Vector for e-commerce product data. For the first The initial weight vector of random projections of hash functions. The first [item] obtained through adaptive learning based on historical access hotspot features of e-commerce product data. Adjust the weight vector for each hot spot. For the first The frequency of access to e-commerce product data This is the adjustment coefficient for the hash weight adjustment based on the access popularity of the e-commerce platform. For the first The offset value of each hash function;

[0148] S43. Based on the collection of e-commerce commodity data storage areas Calculate each storage region The characteristic mean and covariance matrix of the data:

[0149] ;

[0150] ;

[0151] in, Indicates storage area The average feature vector of domestic e-commerce product data, For storage area The covariance matrix of domestic e-commerce product data For storage area The quantity of e-commerce product data within the country, where T is the transpose;

[0152] Furthermore, it optimizes the accuracy of approximate retrieval of e-commerce product data by defining an adaptive hash distance that combines differences in e-commerce product data types and storage area access priorities:

[0153] ;

[0154] in, For the first The and the first The hash distance of e-commerce product data , E-commerce product data and The corresponding number Each hash index value To combine the differences in e-commerce product data types with the first storage areas The access priority is adaptively determined by the first Weight adjustment factor for each hash index bit:

[0155] ;

[0156] in, Adjusting the weights for hash distance calculation based on storage region characteristics. and Storage areas Inner The mean and variance of the dimensional features;

[0157] S44. Based on the hash index vector set and adaptive hash distance constructed in steps S42 and S43, construct a multi-level locality-sensitive hash index table suitable for e-commerce product data location and approximate retrieval:

[0158] ;

[0159] in, For the first A multi-level locality-sensitive hash index table for e-commerce product data in each storage area. For the first A hash index vector of e-commerce product data. For e-commerce product data in the storage area The storage address in This refers to the hash index hierarchy that is divided within the storage area based on the access frequency and data type of e-commerce product data.

[0160] This implementation optimizes e-commerce product data indexing using Locality Sensitive Hash (LSH) technology, achieving efficient low-dimensional mapping of high-dimensional product data, improving data retrieval speed and accuracy. It employs an adaptive hash weight adjustment strategy based on access frequency to ensure index updates adapt to the access characteristics of e-commerce product data, reducing query overhead for high-frequency data. Furthermore, it introduces a multi-level LSH index table structure, hierarchically indexing data with different access priorities, improving index matching accuracy and computational efficiency, and ensuring the high efficiency of the data storage system in high-concurrency query scenarios.

[0161] In this embodiment, S5 includes the following steps:

[0162] S51. E-commerce Commodity Data Storage Solution With multi-level locality sensitive hash index table Establish a global mapping relationship for the e-commerce information storage management system, and define a global data mapping relationship set:

[0163] ;

[0164] in, This refers to the set of global data mapping relationships within the e-commerce information storage and management system. This is the hash index vector corresponding to the e-commerce product data. This is the physical storage address for the e-commerce product data. This refers to the hierarchical information of the e-commerce product data in the multi-level locality-sensitive hash index table;

[0165] S52. Based on the dynamic access frequency of e-commerce product data, dynamically adjust the data storage location of e-commerce product data and define storage migration optimization rules:

[0166] ;

[0167] in, The adjusted location for storing e-commerce product data. To store migration weight parameters, For the first Real-time access priority for e-commerce product data. For storage area The historical average access priority;

[0168] S53. Adjust the hash index table according to changes in access patterns of e-commerce product data. Define the dynamic update rules for the hash index:

[0169] ;

[0170] in, For the first The updated hash index table at each time point. This indicates new index entries added due to new data or data storage migration. These are index deletions caused by data migration or deletion;

[0171] S54. Construct an access path optimization model for e-commerce product data by integrating the global storage management system for e-commerce product data. The e-commerce product data access path optimization model considers data storage location. Hash index table after index update Based on the combined effects, we define the objective function for optimizing e-commerce product data access:

[0172] ;

[0173] in, Optimize the objective function for accessing e-commerce product data. For the first Query response time for e-commerce product data. For the first The computational complexity of querying e-commerce product data. The extent to which index updates affect the query path:

[0174] ;

[0175] in, Adjust the weighting coefficients of the index's impact on the access optimization objective function;

[0176] The impact of index updates on query paths This reflects the impact of index updates on the data access path. The larger the value, the more drastic the index update, which will cause changes in the data query path.

[0177] S55. Based on the optimized data storage scheme Optimization objective function for e-commerce product data access Construct the final e-commerce information storage management system and perform iterative optimization of the storage structure:

[0178] ;

[0179] in, For the first The data of the first e-commerce product in the first The final storage location after the next iteration Adjust the step size factor for global storage. For the first The data of the first e-commerce product in the first The access optimization target value at the next iteration. and The first The minimum and maximum values ​​of the objective function for optimizing e-commerce product data access in each iteration. For the latest storage location, This is the final storage location from the previous round of optimization. Adjusting the impact factor of index updates on storage scheme updates, and considering the degree of impact of index updates on query paths. If the value exceeds the preset value, the storage location will be adjusted to adapt to the index changes and optimize retrieval efficiency.

[0180] This implementation constructs a unified e-commerce information storage management system, enabling synergistic optimization of storage schemes and data index structures to improve overall system performance. It employs a data access optimization objective function, combined with dynamic adjustment of storage location and index update optimization, to improve the data query speed and system response efficiency of the e-commerce platform. By dynamically adjusting storage location through index update influencing factors, it optimizes query paths, improves retrieval accuracy, and ensures that the data storage and indexing system remains in an optimal state in a constantly changing data environment, achieving high efficiency and stability in e-commerce data management.

[0181] An e-commerce information storage system based on big data analytics, used to execute an e-commerce information storage method based on big data analytics, includes the following modules:

[0182] The data preprocessing module is used to receive e-commerce product data from the e-commerce platform, and to standardize the data to generate a standardized e-commerce product dataset.

[0183] The hierarchical storage management module is used to build a hierarchical storage structure based on standardized e-commerce product datasets. According to the access frequency, data type and storage resource allocation requirements of e-commerce product data, it divides e-commerce product data into different storage areas and sets the access priority of each storage area.

[0184] An improved fruit fly optimization storage module is used to optimize the distribution of e-commerce product data in various storage areas and the resource configuration of storage nodes based on the improved fruit fly optimization algorithm, and to establish an optimized e-commerce product data storage scheme.

[0185] The index building and optimization module is used to build an e-commerce product data index on a standardized e-commerce product dataset using locality-sensitive hashing (LSH) technology. It generates a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data and establishes a multi-level LSH index table for the e-commerce product data.

[0186] The information storage management module is used to combine the optimized e-commerce commodity data storage scheme with the multi-level local sensitive hash index table of e-commerce commodity data to build a unified e-commerce information storage management system.

[0187] Example 1:

[0188] On January 10, 2024, during a Chinese New Year shopping festival promotion on a major e-commerce platform, the technical team noticed a significant increase in query response time during peak hours. Some query requests even took more than 2 seconds to respond, resulting in a decline in user search experience, slow page loading, and a surge in complaints. Through log analysis, the system maintenance team discovered that the database access bottleneck was mainly in the query module of the product details page. This module's queries involve a large amount of high-dimensional data, including basic product information, user browsing history, and product ratings. Furthermore, the database storage resource utilization was uneven, with some frequently accessed data stored in inefficient storage areas, affecting query efficiency.

[0189] To verify the feasibility of the method of the present invention, the technical team decided to use the method of the present invention to optimize the storage and accelerate the query of e-commerce product data. The optimization scheme is divided into two stages: the first stage uses the improved fruit fly optimization algorithm for storage optimization, and the second stage uses local sensitive hash index to optimize the query structure.

[0190] At 1:00 AM on January 11, 2024, in order to minimize the impact on user queries, the technical team optimized data storage on the database server. First, they conducted statistical analysis on historical transaction data from January 2023 to January 2024 and found that 80% of query requests were concentrated on product data from the most recent three months, while more than 60% of historical data (such as transaction records from two years ago) was hardly accessed. However, this data was still stored on high-performance storage media, consuming a lot of resources.

[0191] The data on product access frequency from December 2023 to January 2024 was statistically analyzed and divided into high-frequency access data (top 15%), medium-frequency access data (30%), and low-frequency access data (55%).

[0192] High-frequency data is stored in the NVMe SSD storage area, mid-frequency data is stored in the SATA SSD storage area, and low-frequency data is stored in the HDD storage area.

[0193] The dynamic adjustment strategy using the fruit fly optimization algorithm ensures that when the number of visits to a product increases by more than 50% within 24 hours, the data storage location of that product will be automatically adjusted to a higher priority storage area.

[0194] On January 12, 2024, after storage optimization, the query response time of the product details page decreased significantly.

[0195] Taking the "New Year Special Offer Lucky Bag" as an example, the average response time for querying this product on January 10, 2024 was 1.8 seconds. After optimization, when querying the same product on January 12, 2024, the response time was shortened to 0.9 seconds, an optimization of 50%.

[0196] The average response time for low-frequency data access was reduced from 1.5 seconds to 1.1 seconds, resulting in an overall improvement in data access efficiency.

[0197] On January 15, 2024, the technical team further optimized the index structure for product queries to solve the problem of slow retrieval of high-dimensional feature data. For example, during the "Spring Festival promotion", when users search for "New Year's gift boxes", the platform needs to match user purchase preferences, product ratings and historical transaction records. The original index query method took a long time and affected the user experience.

[0198] At 22:00 on January 15, 2024, the technical team began to reconstruct the index of the product database.

[0199] The product index is constructed using locality-sensitive hashing (LSH) technology, and the hash index weight is adjusted based on the product access frequency to increase the index priority of frequently queried products and reduce the amount of query computation.

[0200] For example, products that have been queried more than 10,000 times in the past three days will have their index weight automatically increased, causing queries to be matched in the front layer of the hash index table first, thus reducing query time.

[0201] On January 16, 2024, when querying "Spring Festival Gift Package", the system response time was reduced from 700ms to 410ms, and the query efficiency was improved by 41.4%.

[0202] Taking the query log data from 20:00 on January 16, 2024 as an example:

[0203] Traditional MongoDB B+ tree index query time: average 920ms.

[0204] After optimization using the LSH index of this invention, the average query time is 520ms.

[0205] When searching for "Chinese New Year gift box", the success rate of the query was 92.4% on January 14, 2024 (some queries timed out). After optimization, the success rate increased to 99.1% on January 16, 2024, and the accuracy of the query was improved.

[0206] On January 18, 2024, the e-commerce platform's technical team monitored a surge in queries for a popular product, "Spring Festival Gift Box A," between 3:00 AM and 5:00 AM, with a total of 12,000 queries, an 85% increase compared to 24 hours prior. At 4:30 AM, the Fruit Fly optimization algorithm automatically identified the surge in query popularity for this product and triggered a storage migration mechanism, migrating its data from SATA SSD to NVMe SSD to improve query priority.

[0207] At 12:00 noon on January 18, 2024, the operations team discovered that the click-through rate of the product on the homepage recommendation slot had increased from 2.8% to 5.3%, and the average user dwell time had increased by 30%, directly driving sales growth.

[0208] Experiments conducted from January 10th to January 20th, 2024, verified the effectiveness of the method of this invention on a real e-commerce platform, and compared it with traditional methods. The experimental data are as follows:

[0209]

[0210] The method of this invention optimizes the data storage structure by improving the fruit fly optimization algorithm, so that high-frequency data is stored on high-efficiency storage media, thereby improving query response speed; it optimizes the query structure by combining locality-sensitive hash index, making high-dimensional data retrieval more efficient; and it automatically adjusts the data storage location through a dynamic storage migration strategy to improve storage resource utilization. Experimental results show that the method of this invention has significant improvements in query response speed, storage resource utilization, and system stability, and has high practical value and market application prospects.

[0211] This invention employs an improved fruit fly optimization algorithm to globally optimize and locally tune the storage distribution of e-commerce product data. By combining the access priority of e-commerce product data, adaptive olfactory guidance strategy, and historical access hotspot information, it achieves the optimal distribution of e-commerce product data in the storage area. Furthermore, it introduces an access frequency sensitive factor and a historical hotspot enhancement strategy, enabling the fruit fly optimization algorithm to dynamically adapt to different access patterns of e-commerce product data during data distribution adjustment, thereby reducing the load imbalance of storage nodes and improving the concurrent response capability of data access.

[0212] This invention introduces locality-sensitive hashing technology and combines it with an adaptive weight adjustment strategy based on the access popularity of e-commerce product data to construct a multi-level hash index table. By adaptively adjusting the weight of the hash index, hot data can be preferentially mapped to high-priority storage areas, thereby improving query efficiency.

[0213] This invention constructs an adaptive hierarchical storage structure based on a standardized e-commerce product dataset, and adjusts the data storage location according to the real-time access pattern of e-commerce data through a dynamic storage migration optimization strategy, thereby achieving a reasonable allocation of storage resources. Specifically, the storage migration optimization rule based on access priority enables e-commerce product data to be dynamically adjusted between cold and hot data storage areas, avoiding hot data storage in inefficient storage areas and improving access performance.

[0214] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method and system for storing e-commerce information based on big data analytics, characterized in that, Includes the following steps: S1. Format the e-commerce product data collected from the e-commerce platform to form a standardized e-commerce product dataset; S2. Construct a hierarchical storage structure based on the standardized e-commerce product dataset, divide the e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of the e-commerce product data, and preset the access priority of each storage area; S3. The improved fruit fly optimization algorithm is used to perform global search and local optimization on the hierarchical storage structure to optimize the distribution of e-commerce commodity data in each storage area and the resource allocation of storage nodes, thus forming an optimized e-commerce commodity data storage scheme. S3 includes the following steps: S31. Based on the hierarchical storage structure of e-commerce product data, and with the goals of reducing e-commerce platform response latency, reducing data redundancy, and improving load balancing of storage nodes, establish an objective function for optimizing e-commerce product data storage: ; in, Optimize the objective function for e-commerce product data storage. For the first storage areas The average access latency for e-commerce product data in China This represents the maximum latency for accessing e-commerce product data in each storage region. For the first The proportion of redundant data in e-commerce product data in each storage area This represents the maximum redundancy ratio of e-commerce product data in each storage region. For the first The degree of load imbalance of storage nodes in each storage region This represents the maximum value of the load imbalance between storage nodes in each storage region. , , These represent the optimization weighting coefficients for access latency, data redundancy, and load balancing in e-commerce platforms. Indicates the total number of storage regions; S32. Initialize the population of the improved fruit fly optimization algorithm suitable for optimizing the data storage structure of e-commerce products, and define the initial fruit fly individual position as follows: ; in, This is the initial set of individual locations in the fruit fly population. For the first One fruit fly individual, For population size, Indicates the first Distribution scheme of individual e-commerce product data across various storage areas Indicates the first Initial configuration scheme for storage node resources corresponding to each individual; S33. Combining an adaptive sniffing guidance strategy for prioritizing access to e-commerce product data storage, with the objective function of e-commerce product data storage optimization. To dynamically update the location of individual fruit flies: ; in, For the next iteration Individual fruit fly location update scheme The global optimization sensitivity coefficient, For access priority sensitive factors, For the first Storage area access priority, This represents the optimal storage location scheme in the current iteration. This is the random disturbance adjustment coefficient. To combine storage areas Lévy flight random step size for access priority adjustment For the first The objective function value for optimizing e-commerce product data storage for each individual entity. Optimize the objective function value for storing the optimal e-commerce product data in the current population; S34. Based on historical access hotspot information from the e-commerce platform, construct a memory-enhanced local clustering optimization strategy to further update the individual fruit fly positions: ; in, A memory-enhanced e-commerce product data storage location scheme. Sensitivity coefficient for historical hot topics Indicates the first Each individual corresponds to a storage area Historical access frequency, For the first Historical average access frequency of each storage region For storage area The local optimal storage location is determined by the access frequency of historical hot e-commerce product data. S35. Repeat steps S33 and S34 above, and dynamically adjust the global optimization sensitivity coefficient according to the real-time adaptability changes of the e-commerce commodity data storage structure. Random disturbance adjustment coefficient Sensitivity coefficient to historical hot topics The scheme for the location of individual fruit fly populations is continuously iterated and updated until the objective function for optimizing e-commerce commodity data storage is optimized. Convergence, output the final optimized e-commerce product data storage solution: ; in, This represents the final optimized scheme for the location of individual fruit fly individuals in the population. To optimize the storage distribution of e-commerce product data across various storage areas, The optimized storage node resource configuration scheme; S4. The standardized e-commerce product dataset is constructed using locality-sensitive hashing (LSH) technology to create an e-commerce product data index. A low-dimensional hash index is generated based on the high-dimensional feature information of the e-commerce product data, and a unified multi-level LSH index table for e-commerce product data is established in each storage area. S5. The optimized e-commerce commodity data storage scheme is combined with a multi-level local sensitive hash index table to form a unified e-commerce information storage management system.

2. The e-commerce information storage method and system based on big data analysis according to claim 1, characterized in that, S1 includes the following steps: S11. Collect e-commerce product data from the e-commerce platform, wherein the e-commerce product data includes basic product information, classification information, historical transaction records and user interaction information, and construct an e-commerce product dataset; S12. Format the e-commerce product dataset so that all data items in the e-commerce product dataset meet a unified data storage format standard; S13. Perform a consistency check on the formatted e-commerce product dataset to ensure that all data items within the dataset meet the requirements of integrity, correctness, and uniqueness. Define the formatted standardized e-commerce product dataset as follows: ; in, For the standardized e-commerce product dataset, For the formatted first Each product data record , , , These respectively represent standardized basic product information, product category information, historical transaction records, and user interaction information. This represents the total amount of product data.

3. The e-commerce information storage method and system based on big data analysis according to claim 2, characterized in that, S2 includes the following steps: S21. Based on the standardized e-commerce product dataset To analyze the access frequency, data type, and storage resource requirements of e-commerce product data, an e-commerce product data feature matrix is ​​established. ; in, For e-commerce product data feature matrix, Indicates the first A set of features for e-commerce product data records. For the first The frequency of access to e-commerce product data For the first Data types of e-commerce product data For the first Storage resource requirements for e-commerce product data; S22. Based on the e-commerce product data feature matrix Determine the storage area partitioning rules for e-commerce product data, and map the e-commerce product data to different storage areas: ; in, This represents a collection of data storage areas for e-commerce product information. For the first One storage area, Indicates the total number of storage areas. , , They represent the first The access frequency range, data type range, and storage resource allocation rules corresponding to each storage region; S23. For each storage region Preset access priorities, defining the access priorities for each storage area as follows: ; in, Indicates the first Access priority of each storage region , , The first Normalized average values ​​of e-commerce product data access frequency, data type, and storage resource requirements within each storage region. , , These are the weighting coefficients for access frequency, data type, and storage resource requirements, respectively. S24. Based on storage region access priority For the set of storage regions Sort the data to complete the hierarchical storage structure of e-commerce product data.

4. The e-commerce information storage method and system based on big data analysis according to claim 3, characterized in that, S4 includes the following steps: S41. Based on standardized e-commerce product datasets Extract high-dimensional feature information from e-commerce product data and construct a high-dimensional feature matrix suitable for indexing e-commerce product data: ; in, This represents a high-dimensional feature matrix of e-commerce product data. For the first A high-dimensional feature vector of e-commerce product data Indicates the first The first e-commerce product data 3D eigenvalues The total dimension of the high-dimensional features. The total amount of e-commerce product data; S42. Based on the historical access frequency characteristics of e-commerce platform product data, an adaptive hash weight adjustment strategy for e-commerce product data access popularity is introduced. A method for generating locally sensitive hash index vectors for e-commerce product data that adapts to access popularity is constructed, defining the... The hash index vector of each e-commerce product data is: ; in, For the first A low-dimensional hash index vector of e-commerce product data. For the first The first e-commerce product data There are several hash index values, where L is the dimension of the Local Sensitive Hash Index Vector for e-commerce product data. For the first The initial weight vector of random projections of hash functions. The first [item] obtained through adaptive learning based on historical access hotspot features of e-commerce product data. Adjust the weight vector for each hot spot. For the first The frequency of access to e-commerce product data This is the adjustment coefficient for the hash weight adjustment based on the access popularity of the e-commerce platform. For the first The offset value of each hash function; S43. Based on the collection of e-commerce commodity data storage areas Calculate each storage region The characteristic mean and covariance matrix of the data: ; ; in, Indicates storage area The average feature vector of domestic e-commerce product data, For storage area The covariance matrix of domestic e-commerce product data For storage area The quantity of e-commerce product data within the country, where T is the transpose; Furthermore, it optimizes the accuracy of approximate retrieval of e-commerce product data by defining an adaptive hash distance that combines differences in e-commerce product data types and storage area access priorities: ; in, For the first The and the first The hash distance of e-commerce product data , E-commerce product data and The corresponding number Each hash index value To combine the differences in e-commerce product data types with the first storage areas The access priority is adaptively determined by the first Weight adjustment factor for each hash index bit: ; in, Adjusting the weights for hash distance calculation based on storage region characteristics. and Storage areas Inner The mean and variance of the dimensional features; S44. Based on the hash index vector set and adaptive hash distance constructed in steps S42 and S43, construct a multi-level locality-sensitive hash index table suitable for e-commerce product data location and approximate retrieval: ; in, For the first A multi-level locality-sensitive hash index table for e-commerce product data in each storage area. For the first A hash index vector of e-commerce product data. For e-commerce product data in the storage area The storage address in This refers to the hash index hierarchy that is divided within the storage area based on the access frequency and data type of e-commerce product data.

5. The e-commerce information storage method and system based on big data analysis according to claim 4, characterized in that, S5 includes the following steps: S51. E-commerce Commodity Data Storage Solution With multi-level locality sensitive hash index table Establish a global mapping relationship for the e-commerce information storage management system, and define a global data mapping relationship set: ; in, This refers to the set of global data mapping relationships within the e-commerce information storage and management system. This is the hash index vector corresponding to the e-commerce product data. This is the physical storage address for the e-commerce product data. This refers to the hierarchical information of the e-commerce product data in the multi-level locality-sensitive hash index table; S52. Based on the dynamic access frequency of e-commerce product data, dynamically adjust the data storage location of e-commerce product data and define storage migration optimization rules: ; in, The adjusted location for storing e-commerce product data. To store migration weight parameters, For the first Real-time access priority for e-commerce product data. For storage area The historical average access priority; S53. Adjust the hash index table according to changes in access patterns of e-commerce product data. Define the dynamic update rules for the hash index: ; in, For the first The updated hash index table at each time point. This indicates new index entries added due to new data or data storage migration. These are index deletions caused by data migration or deletion; S54. Construct an access path optimization model for e-commerce product data by integrating the global storage management system for e-commerce product data. The e-commerce product data access path optimization model considers data storage location. Hash index table after index update Based on the combined effects, we define the objective function for optimizing e-commerce product data access: ; in, Optimize the objective function for accessing e-commerce product data. For the first Query response time for e-commerce product data. For the first The computational complexity of querying e-commerce product data. The extent to which index updates affect the query path: ; in, Adjust the weighting coefficients of the index's impact on the access optimization objective function; The impact of index updates on query paths This reflects the impact of index updates on the data access path. The larger the value, the more drastic the index update, which will cause changes in the data query path. S55. Based on the optimized data storage scheme Optimization objective function for e-commerce product data access Construct the final e-commerce information storage management system and perform iterative optimization of the storage structure: ; in, For the first The data of the first e-commerce product in the first The final storage location after the next iteration Adjust the step size factor for global storage. For the first The data of the first e-commerce product in the first The access optimization target value at the next iteration. and The first The minimum and maximum values ​​of the objective function for optimizing e-commerce product data access in each iteration. For the latest storage location, This is the final storage location from the previous round of optimization. Adjusting the impact factor of index updates on storage scheme updates, and considering the degree of impact of index updates on query paths. If the value exceeds the preset value, the storage location will be adjusted to adapt to the index changes and optimize retrieval efficiency.

6. An e-commerce information storage system based on big data analysis, used to execute the e-commerce information storage method based on big data analysis as described in any one of claims 1-5, characterized in that, Includes the following modules: The data preprocessing module is used to receive e-commerce product data from the e-commerce platform, and to standardize the data to generate a standardized e-commerce product dataset. The hierarchical storage management module is used to build a hierarchical storage structure based on standardized e-commerce product datasets. According to the access frequency, data type and storage resource allocation requirements of e-commerce product data, it divides e-commerce product data into different storage areas and sets the access priority of each storage area. An improved fruit fly optimization storage module is used to optimize the distribution of e-commerce product data in various storage areas and the resource configuration of storage nodes based on the improved fruit fly optimization algorithm, and to establish an optimized e-commerce product data storage scheme. The index building and optimization module is used to build an e-commerce product data index on a standardized e-commerce product dataset using locality-sensitive hashing (LSH) technology. It generates a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data and establishes a multi-level LSH index table for the e-commerce product data. The information storage management module is used to combine the optimized e-commerce commodity data storage scheme with the multi-level local sensitive hash index table of e-commerce commodity data to build a unified e-commerce information storage management system.

Citation Information

Patent Citations

  • Edge data storage method based on aging intimacy model

    CN115865712A

  • Reservoir area layout method and device based on category association rule and storage medium

    CN117333113A