Electronic commerce information storage method and system based on big data analysis

By improving the fruit fly optimization algorithm and locally sensitive hashing technology, dynamically adjusting the storage location and index structure of e-commerce data, the problem of inefficient data storage on e-commerce platforms is solved, and efficient query and stable storage are achieved.

CN120277068AActive Publication Date: 2025-07-08BEIJING ZHISUANDUODUO TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510333523.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing e-commerce data storage solutions have degraded query performance in massive data environments, high index maintenance overhead, poor scalability, rough hot and cold data management, and failed to optimize in combination with access mode, resulting in low storage efficiency.

Method used

The improved fruit fly optimization algorithm is used for global optimization and local tuning, and a multi-level hash index table is built with local sensitive hash technology, dynamically adjusting the data storage location, and optimizing storage resource allocation and query efficiency.

Benefits of technology

It improves data access performance, reduces load imbalance of storage nodes, improves query efficiency and storage resource utilization, and ensures the stability and efficiency of the system in high concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277068A_ABST
    Figure CN120277068A_ABST
Patent Text Reader

Abstract

The invention discloses an e-commerce information storage method and system based on big data analysis. The method comprises the following steps: S1, formatting e-commerce commodity data collected in an e-commerce platform to form a standardized e-commerce commodity data set; s2, dividing the e-commerce commodity data into different storage areas according to the access frequency, the data type and the storage resource allocation requirement of the e-commerce commodity data; s3, forming an optimized e-commerce commodity data storage scheme; s4, establishing a unified e-commerce commodity data multi-level locality sensitive hash index table in each storage area; and S5, combining the optimized e-commerce commodity data storage scheme with the e-commerce commodity data index system to form a unified e-commerce information storage management system. According to the invention, electronic commerce commodity data can be dynamically adjusted between cold and hot data storage areas, hot data is prevented from being stored in a low-efficiency storage area, and the access performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information storage, and particularly to an e-commerce information storage method and system based on big data analysis. Background Art

[0002] With the rapid development of the e-commerce industry, the commodity data on e-commerce platforms has shown explosive growth. The multi-dimensional information such as the types of commodities, transaction records, and user interaction data has been continuously accumulating. The storage and management of such large-scale data have become one of the key issues in the operation of e-commerce platforms. The storage of e-commerce information not only requires efficient access capabilities to ensure fast query and response of users, but also needs to optimize the data storage structure to reduce storage costs and improve the scalability of data management. In the existing technologies, the storage of e-commerce commodity data mainly adopts traditional relational databases or distributed storage solutions, but these solutions have obvious limitations.

[0003] Currently, traditional e-commerce data storage methods usually rely on relational databases or NoSQL databases. Although relational databases have a mature management mechanism in structured data storage, with the growth of e-commerce data volume, their query performance gradually decreases. Especially in the environment of massive data, the overhead of index maintenance is too high, affecting the response speed of the system. In addition, the scalability of relational databases is poor and it is difficult to support the needs of e-commerce platforms for large-scale concurrent access.

[0004] In order to alleviate the storage and query performance bottlenecks of relational databases, in recent years, many e-commerce platforms have introduced NoSQL databases and distributed storage systems. However, since e-commerce data usually contains a large amount of historical transaction records and high-dimensional data such as user interaction information, NoSQL databases still have deficiencies in handling complex queries and transaction consistency. In addition, most of the existing NoSQL databases lack an efficient index optimization mechanism, resulting in the query efficiency still being difficult to meet the high-performance requirements in a large-scale data environment.

[0005] In addition, with the diverse requirements for data storage on e-commerce platforms, the hierarchical management of hot and cold data access has become one of the key technologies for optimizing storage. Some current data storage optimization schemes mainly divide data into cold data and hot data based on the access frequency or time characteristics of the data, and use different storage media for hierarchical management. However, the simple method of dividing hot and cold data has the following problems: it cannot make full use of the type characteristics of commodity data and the storage resource allocation requirements, resulting in a lack of targeted optimization of the storage structure; the data migration strategy is relatively crude and fails to carry out refined management in combination with the dynamic access mode, affecting the data access efficiency.

[0006] In addition, there are still certain deficiencies in the existing e-commerce data storage optimization methods in terms of global optimization and local tuning. For example, some systems use heuristic algorithms or traditional optimization methods to adjust the data storage structure. However, due to the failure to fully combine the access priorities, storage resource requirements, and load balancing characteristics of e-commerce data, it is often difficult to balance the optimization of storage performance and the efficiency of data access.

[0007] In summary, the existing technologies mainly have the following problems in the storage of e-commerce commodity data: First, the traditional database solution is difficult to meet the high-performance requirements of massive data storage and query; Second, the existing NoSQL storage solutions lack an efficient index optimization mechanism, resulting in limited query efficiency; Third, the existing cold and hot data management methods are relatively rough and fail to perform intelligent optimization by combining data characteristics and access patterns; Fourth, the existing storage optimization algorithms have deficiencies in global and local tuning and are difficult to achieve efficient data distribution management. Therefore, there is an urgent need for an e-commerce information storage method based on big data analysis to improve the flexibility, query efficiency, and optimization effect of data storage. Summary of the Invention

[0008] An object of the present invention is to provide an e-commerce information storage method and system based on big data analysis, which enables the dynamic adjustment of e-commerce commodity data between cold and hot data storage areas, avoids storing hot data in inefficient storage areas, and improves the access performance.

[0009] An e-commerce information storage method and system based on big data analysis according to an embodiment of the present invention includes the following steps:

[0010] S1. Format the e-commerce commodity data collected from the e-commerce platform to form a standardized e-commerce commodity data set;

[0011] S2. Based on the standardized e-commerce commodity data set, construct a hierarchical storage structure, divide the e-commerce commodity data into different storage areas according to the access frequency, data type, and storage resource allocation requirements of the e-commerce commodity data, and preset the access priorities of each storage area;

[0012] S3. Use the improved fruit fly optimization algorithm to perform global search and local tuning on the hierarchical storage structure, optimize the distribution of e-commerce commodity data in each storage area and the resource configuration of storage nodes, and form an optimized e-commerce commodity data storage solution;

[0013] S4. Apply the locality-sensitive hashing technique to the standardized e-commerce product data set to construct an e-commerce product data index, generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data, and establish a unified multi-level locality-sensitive hash index table for e-commerce product data in each storage area;

[0014] S5. Combine the optimized e-commerce product data storage scheme with the multi-level locality-sensitive hash index table to form a unified e-commerce information storage and management system.

[0015] Optionally, S1 includes the following steps:

[0016] S11. Collect e-commerce product data in the e-commerce platform, where the e-commerce product data includes product basic information, classification information, historical transaction records, and user interaction information, and construct an e-commerce product data set;

[0017] S12. Perform formatting processing on the e-commerce product data set so that all data items in the e-commerce product data set meet the unified data storage format standard;

[0018] S13. Perform consistency checking on the formatted e-commerce product data set so that all data items in the e-commerce product data set meet the requirements of integrity, correctness, and uniqueness, and define the formatted standardized e-commerce product data set as:

[0019]

[0020] Among them, D std is the standardized e-commerce product data set, d′ i is the i-th product data record after formatting, I′ i , C′ i , T′ i , U′ i respectively represent the standardized product basic information, product classification information, historical transaction records, and user interaction information, and N is the total amount of product data.

[0021] Optionally, S2 includes the following steps:

[0022] S21. Based on the standardized e-commerce product data set D std , count the access frequency, data type, and storage resource requirements of the e-commerce product data, and establish an e-commerce product data feature matrix:

[0023]

[0024] Among them, F matrix is the e-commerce product data feature matrix, fi Denote the feature set of the \(i\)-th e-commerce product data record as \(F\). i Let \(Type\) be the access frequency of the \(i\)-th e-commerce product data. i Let \(Res\) be the data type of the \(i\)-th e-commerce product data. i Let \(Res\) be the storage resource requirement of the \(i\)-th e-commerce product data.

[0025] S22. Determine the storage area division rule for the e-commerce product data according to the e-commerce product data feature matrix \(F\), and map the e-commerce product data to different storage areas: matrix where \(Z\) represents the set of e-commerce product data storage areas, \(Z_j\) is the \(j\)-th storage area, \(M\) represents the total number of storage areas, \([\![Type_{jmin}, Type_{jmax}]\!]\), \([\![Res_{jmin}, Res_{jmax}]\!]\), and \([\![Freq_{jmin}, Freq_{jmax}]\!]\) respectively represent the access frequency range, data type range, and storage resource allocation rule corresponding to the \(j\)-th storage area;

[0026]

[0027] where \(Z\) represents the set of e-commerce product data storage areas, \(Z\) j is the \(j\)-th storage area, \(M\) represents the total number of storage areas, respectively represent the access frequency range, data type range, and storage resource allocation rule corresponding to the \(j\)-th storage area;

[0028] S23. Preset an access priority for each storage area \(Z_j\), and define the access priority of each storage area as: j where \([\![Priority_j]\!]\) represents the access priority of the \(j\)-th storage area, \([\![\overline{Type_j}]\!]\), \([\![\overline{Res_j}]\!]\), and \([\![\overline{Freq_j}]\!]\) are the normalized average values of the access frequency, data type, and storage resource requirement of the e-commerce product data in the \(j\)-th storage area respectively, and \(\omega_1\), \(\omega_2\), and \(\omega_3\) are the weight coefficients of the access frequency, data type, and storage resource requirement respectively;

[0029]

[0030] where, represents the access priority of the \(j\)-th storage area, are the normalized average values of the access frequency, data type, and storage resource requirement of the e-commerce product data in the \(j\)-th storage area respectively, and \(\omega_1\), \(\omega_2\), \(\omega_3\) are the weight coefficients of the access frequency, data type, and storage resource requirement respectively;

[0031] S24. Sort the set of storage areas \(Z\) according to the storage area access priority \([\![Priority_j]\!]\) to complete the construction of the hierarchical storage structure of the e-commerce product data.

[0032] Optionally, the step S3 includes the following steps:

[0033] S31. According to the hierarchical storage structure of the e-commerce product data, with the goal of reducing the response latency of the e-commerce platform, reducing data redundancy, and improving the load balance of the storage nodes, establish an optimization objective function for the storage of the e-commerce product data:

[0034]

[0035] where \(G\) opt is the optimization objective function for the storage of the e-commerce product data,​ is the average access latency of e-commerce commodity data in the j-th storage area Z j C max is the maximum value of the access latency of e-commerce commodity data in each storage area is the redundancy data ratio of e-commerce commodity data in the j-th storage area, R max is the maximum value of the redundancy ratio of e-commerce commodity data in each storage area is the imbalance degree of the storage node load in the j-th storage area, L max is the maximum value of the imbalance degree of the storage node load in each storage area. λ1, λ2, and λ3 are the optimization weight coefficients of access latency, data redundancy, and load balancing in the e-commerce platform respectively;

[0036] S32. Initialize the improved fruit fly optimization algorithm population suitable for the optimization of the e-commerce commodity data storage structure, and establish the initial fruit fly individual position definition as:

[0037]

[0038] where X init is the initial fruit fly population individual position set, x k is the k-th fruit fly individual, K is the population size, Pos k represents the distribution scheme of the e-commerce commodity data of the k-th individual in each storage area, Node k represents the initial configuration scheme of the storage node resources corresponding to the k-th individual;

[0039] S33. Combine the adaptive olfactory guidance strategy of the e-commerce commodity data storage access priority, and dynamically update the fruit fly individual position based on the e-commerce commodity data storage optimization objective function G opt as follows:

[0040]

[0041] where is the update scheme of the k-th fruit fly individual position in the next iteration, μ is the global optimization sensitivity coefficient, γ is the access priority sensitivity factor, is the access priority of the j-th storage area defined in claim 3, is the optimal storage position scheme in the current iteration, η is the random perturbation adjustment coefficient, is combined with the storage area Z j the Lévy flight random step size adjusted by the access priority, is the value of the e-commerce commodity data storage optimization objective function corresponding to the k-th individual, is the optimal value of the e-commerce commodity data storage optimization objective function of the current population;

[0042] S34. Construct a memory-enhanced local aggregation optimization strategy based on the historical access hot-spot information of the e-commerce platform, and further update the positions of the fruit fly individuals:

[0043]

[0044] wherein, is the storage location scheme of the e-commerce commodity data after memory enhancement, θ is the historical hot-spot sensitivity coefficient, represents the historical access frequency of the k-th individual corresponding to the storage area Z j , is the historical average access frequency of the j-th storage area, is the local optimal storage location determined based on the historical hot-spot e-commerce commodity data access frequency within the storage area Z j ;

[0045] S35. Repeat the above steps S33 and S34, dynamically adjust the global optimization sensitivity coefficient μ, the random perturbation adjustment coefficient η, and the historical hot-spot sensitivity coefficient θ according to the real-time fitness change of the e-commerce commodity data storage structure, and continuously iteratively update the individual position scheme of the fruit fly population until the e-commerce commodity data storage optimization objective function G opt converges, and output the finally optimized e-commerce commodity data storage scheme:

[0046] X final ={x best |x best =(Pos best , Node best )};

[0047] wherein, X final represents the finally optimized individual position scheme of the fruit fly population, Pos best is the storage distribution position of the optimized e-commerce commodity data in each storage area, and Node best is the optimized storage node resource allocation scheme.

[0048] Optionally, the S4 includes the following steps:

[0049] S41. Extract the high-dimensional feature information of the e-commerce commodity data based on the standardized e-commerce commodity data set D std , and construct a high-dimensional feature matrix suitable for indexing the e-commerce commodity data:

[0050]

[0051] wherein, V high represents the high-dimensional feature matrix of the e-commerce commodity data, v iis the high-dimensional feature vector of the i-th e-commerce product data, v id represents the d-th feature value of the i-th e-commerce product data, D is the total dimension of the high-dimensional features, and N is the total amount of e-commerce product data;

[0052] S42. According to the historical access frequency characteristics of the e-commerce platform product data, introduce an adaptive hash weight adjustment strategy for the access heat of e-commerce product data, construct a method for generating a locality-sensitive hash index vector of e-commerce product data with adaptive access heat, and define the hash index vector of the i-th e-commerce product data as:

[0053]

[0054] where, H i is the low-dimensional hash index vector of the i-th e-commerce product data, h il is the l-th hash index value of the i-th e-commerce product data, L is the dimension of the locality-sensitive hash index vector of e-commerce product data, W l is the initial weight vector of the random projection of the l-th hash function, is the l-th hot spot adjustment weight vector adaptively learned based on the historical access hot spot characteristics of e-commerce product data, F i is the access frequency of the i-th e-commerce product data, ρ is the adjustment coefficient of the e-commerce platform access heat on the hash weight adjustment, b l is the offset value of the l-th hash function;

[0055] S43. Based on the set Z of e-commerce product data storage areas, calculate the feature mean and covariance matrix of the data in each storage area Z j :

[0056]

[0057] where, represents the average feature vector of the e-commerce product data in the storage area Z j , is the covariance matrix of the e-commerce product data in the storage area Z j , is the number of e-commerce product data in the storage area Z j , and T is the transpose;

[0058] And optimize the approximate retrieval accuracy of e-commerce product data, and define an adaptive hash distance that combines the e-commerce product data type difference and the storage area access priority:

[0059]

[0060] where, D(Hi , H j ), which is the Hamming distance between the i-th and j-th e-commerce product data, h il , h jl are respectively the l-th hash index values corresponding to the e-commerce product data i and j, is the l-th hash index bit weight adjustment coefficient adaptively determined by combining the e-commerce product data type difference and the access priority of the j-th storage area Z j :

[0061]

[0062] where β is the adjustment weight of the storage area characteristics on the Hamming distance calculation, and are respectively the mean and variance of the l-th dimension characteristics in the storage area Z j ;

[0063] S44. Based on the hash index vector set and the adaptive Hamming distance constructed in steps S42 and S43, construct a multi-level locality-sensitive hash index table suitable for e-commerce product data location and approximate retrieval:

[0064]

[0065] where is the multi-level locality-sensitive hash index table of the e-commerce product data in the j-th storage area, H i is the hash index vector of the i-th e-commerce product data, Addr i is the storage address of the e-commerce product data in the storage area Z j , and Layer i is the hash index level divided in the storage area based on the e-commerce product data access frequency and data type.

[0066] Optionally, the S5 includes the following steps:

[0067] S51. Based on the e-commerce product data storage scheme Pos best and the multi-level locality-sensitive hash index table establish the global mapping relationship of the e-commerce information storage management system, and define the global data mapping relationship set:

[0068] M store = {(v i , Pos best,i , H i , Addr i , Layer i ) | v i ∈ D std};

[0069] Among them, M store is the global data mapping relationship set in the e-commerce information storage management system, and H i is the hash index vector corresponding to the e-commerce product data, and Addr i is the physical storage address of the e-commerce product data, and Layer i is the hierarchical information of the e-commerce product data in the multi-level locality-sensitive hashing index table;

[0070] S52. Dynamically adjust the storage location of e-commerce product data according to the dynamic access frequency of e-commerce product data, and define the storage migration optimization rule:

[0071]

[0072] Among them, Pos new,i is the adjusted storage location of the e-commerce product data, and α new is the storage migration weight parameter, and P access,i is the real-time access priority of the i-th e-commerce product data, is the average value of the historical access priorities of the storage area Z j ;

[0073] S53. Adjust the hashing index table according to the change of the access mode of e-commerce product data Define the hashing index dynamic update rule:

[0074]

[0075] Among them, is the updated hashing index table at the (t + 1)-th moment, represents the newly added index items due to new data or data storage migration, is the index deletion items due to data migration or deletion;

[0076] S54. Combine the global storage management system of e-commerce product data to construct an access path optimization model for e-commerce product data. The access path optimization model of e-commerce product data considers the comprehensive influence of the data storage location Pos new,i and the hashing index table after index update , and define the access optimization objective function of e-commerce product data:

[0077]

[0078] Among them, G access is the access optimization objective function of e-commerce product data, and T query,iis the query response time for the i-th e-commerce product data, C query,i is the computational complexity of the i-th e-commerce product data query, D index,i is the degree of influence of index update on the query path:

[0079]

[0080] where γ1 is the influence weight coefficient of index adjustment on the access optimization objective function;

[0081] The degree of influence D of index update on the query path index,i reflects the impact of changes in index update on the data access path. The larger the value, the more intense the index update, which will cause changes in the data query path;

[0082] S55. According to the optimized data storage scheme Pos new,i and the e-commerce product data access optimization objective function G access Construct the final e-commerce information storage management system and perform storage structure optimization iteration:

[0083]

[0084] where, is the final storage location of the i-th e-commerce product data after the (t + 1)-th iteration, λ1 is the global storage adjustment step size coefficient, is the access optimization objective value of the i-th e-commerce product data at the t-th iteration, and are respectively the minimum and maximum values of the e-commerce product data access optimization objective function in the t-th iteration, Pos new,i is the latest storage location, is the final storage location of the previous round of optimization, ξ is the influence factor of index adjustment on the update of the storage scheme. If the degree of influence of index update on the query path is greater than the preset value, then adjust the storage location to adapt to the change of the index and optimize the retrieval efficiency.

[0085] An e-commerce information storage system based on big data analysis is used to execute an e-commerce information storage method based on big data analysis, including the following modules:

[0086] A storage data preprocessing module is used to receive e-commerce product data in an e-commerce platform and perform standardization processing on the data to generate a standardized e-commerce product data set;

[0087] Hierarchical storage management module, which is used to construct a hierarchical storage structure based on a standardized e-commerce product data set, divide e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of e-commerce product data, and set the access priorities of each storage area;

[0088] Improved fruit fly optimization storage optimization module, which is used to optimize the distribution of e-commerce product data in each storage area and the storage node resource configuration based on the improved fruit fly optimization algorithm, and establish an optimized e-commerce product data storage solution;

[0089] Index construction and optimization module, which is used to construct an e-commerce product data index for the standardized e-commerce product data set by using the locality-sensitive hashing technique, generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data, and establish a multi-level locality-sensitive hashing index table for the e-commerce product data;

[0090] Information storage management module, which is used to construct a unified e-commerce information storage management system by combining the optimized e-commerce product data storage solution and the multi-level locality-sensitive hashing index table of the e-commerce product data.

[0091] The beneficial effects of the present invention are as follows:

[0092] (1) The present invention uses the improved fruit fly optimization algorithm to globally optimize and locally tune the storage distribution of e-commerce product data. By combining the access priority of e-commerce product data, the adaptive olfactory guidance strategy and the historical access hot spot information, the optimal distribution of e-commerce product data in the storage area is realized. In addition, the access frequency sensitive factor and the historical hot spot enhancement strategy are introduced, so that the fruit fly optimization algorithm can dynamically adapt to different access patterns of e-commerce product data during the data distribution adjustment process, thereby reducing the load imbalance of storage nodes and improving the concurrent response ability of data access.

[0093] (2) The present invention introduces the locality-sensitive hashing technique and combines it with the access heat adaptive weight adjustment strategy of e-commerce product data to construct a multi-level hash index table. By adaptively adjusting the weight of the hash index, hot data can be preferentially mapped to the high-priority storage area, thereby improving the query efficiency.

[0094] (3) The present invention constructs an adaptive hierarchical storage structure based on a standardized e-commerce product data set, and through the dynamic storage migration optimization strategy, adjusts the data storage location according to the real-time access pattern of e-commerce data, so as to realize the reasonable allocation of storage resources. Specifically, based on the storage migration optimization rule of access priority, e-commerce product data can be dynamically adjusted between the hot and cold data storage areas, avoiding hot data being stored in inefficient storage areas and improving the access and storage performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0096] Figure 1 is a flowchart of a method and system for storing e-commerce information based on big data analysis proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0097] The present invention will now be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0098] Refer to Figure 1 , a method and system for storing e-commerce information based on big data analysis, including the following steps:

[0099] S1. Format the e-commerce product data collected from the e-commerce platform to form a standardized e-commerce product data set;

[0100] S2. Build a hierarchical storage structure based on the standardized e-commerce product data set, divide the e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of the e-commerce product data, and preset the access priorities of each storage area;

[0101] S3. Use the improved fruit fly optimization algorithm to perform global search and local optimization on the hierarchical storage structure, optimize the distribution of e-commerce product data in each storage area and the resource configuration of storage nodes, and form an optimized storage scheme for e-commerce product data;

[0102] S4. Use the local sensitive hashing technique to build an index for the e-commerce product data set of the standardized e-commerce product data, generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data, and establish a unified multi-level local sensitive hash index table for the e-commerce product data in each storage area;

[0103] S5. Combine the optimized storage scheme for e-commerce product data with the multi-level local sensitive hash index table to form a unified e-commerce information storage management system.

[0104] In this embodiment, S1 includes the following steps:

[0105] S11. Collect e-commerce product data from the e-commerce platform. The e-commerce product data includes product basic information, classification information, historical transaction records and user interaction information, and build an e-commerce product data set;

[0106] S12. Format the e-commerce product dataset so that all data items in the e-commerce product dataset meet the unified data storage format standard;

[0107] S13. Perform a consistency check on the formatted e-commerce product dataset so that all data items in the e-commerce product dataset meet the requirements of integrity, correctness, and uniqueness. Define the formatted standardized e-commerce product dataset as:

[0108]

[0109] where, D std is the standardized e-commerce product dataset, d′ i is the i-th product data record after formatting, I′ i , C′ i , T′ i , U′ i respectively represent the basic product information, product classification information, historical transaction records, and user interaction information after standardization, and N is the total amount of product data.

[0110] In this embodiment, by standardizing the e-commerce product data, the consistency of the data format is ensured, the reliability of data storage and retrieval is improved, a structured data organization method is adopted, so that the product information, transaction records, and user interaction data have a unified storage standard, data redundancy is reduced, and the accuracy of data query is improved. In addition, through formatting processing and consistency verification, the data integrity and correctness are effectively improved, the storage and calculation errors caused by data format mismatch are reduced, and a solid foundation is provided for subsequent data storage optimization and index construction.

[0111] In this embodiment, S2 includes the following steps:

[0112] S21. Based on the standardized e-commerce product dataset D std , count the access frequency, data type, and storage resource requirements of the e-commerce product data, and establish an e-commerce product data feature matrix:

[0113]

[0114] where, F matrix is the e-commerce product data feature matrix, f i represents the feature set of the i-th e-commerce product data record, F i is the access frequency of the i-th e-commerce product data, Type i is the data type of the i-th e-commerce product data, Res iis the storage resource requirement for the i-th e-commerce product data;

[0115] S22. According to the e-commerce product data feature matrix F matrix Determine the storage area division rules for e-commerce product data, and map the e-commerce product data to different storage areas:

[0116]

[0117] where Z represents the set of e-commerce product data storage areas, Z j is the j-th storage area, M represents the total number of storage areas, respectively represent the access frequency range, data type range, and storage resource allocation rule corresponding to the j-th storage area;

[0118] S23. For each storage area Z j Preset the access priority, and define the access priority of each storage area as:

[0119]

[0120] where, represents the access priority of the j-th storage area, are respectively the normalized average values of the access frequency, data type, and storage resource requirements of the e-commerce product data in the j-th storage area, and ω1, ω2, and ω3 are respectively the weight coefficients of the access frequency, data type, and storage resource requirements;

[0121] S24. According to the storage area access priority Sort the storage area set Z to complete the construction of the hierarchical storage structure of e-commerce product data.

[0122] This embodiment constructs a hierarchical storage structure based on the access frequency, data type, and storage resource allocation requirements of e-commerce product data, improving the flexibility of data management and storage efficiency. By presetting the access priority of the storage area, high-frequency access data is stored in a low-latency storage area, optimizing the data retrieval speed. At the same time, adopting a hierarchical data storage method reduces storage resource waste, improves the system load balancing ability, and ensures that the e-commerce platform can still maintain an efficient and stable operation state in a large-scale data storage environment.

[0123] In this embodiment, S3 includes the following steps:

[0124] S31. According to the hierarchical storage structure of e-commerce product data, with the goal of reducing the response latency of the e-commerce platform, reducing data redundancy, and improving the load balancing of storage nodes, establish an optimization objective function for e-commerce product data storage:

[0125]

[0126] Among them, G opt is the optimization objective function for e-commerce commodity data storage, is the average access latency of e-commerce commodity data in the j-th storage area Z j and C max is the maximum value of the access latency of e-commerce commodity data in each storage area. is the redundancy data ratio of e-commerce commodity data in the j-th storage area, and R max is the maximum value of the redundancy ratio of e-commerce commodity data in each storage area. is the imbalance degree of the load of storage nodes in the j-th storage area, and L max is the maximum value of the imbalance degree of the load of storage nodes in each storage area. λ1, λ2, and λ3 are the optimization weight coefficients of access latency, data redundancy, and load balancing in the e-commerce platform respectively;

[0127] S32. Initialize the population of the improved fruit fly optimization algorithm applicable to the optimization of the e-commerce commodity data storage structure, and establish the definition of the initial fruit fly individual position as:

[0128]

[0129] Among them, X init is the set of individual positions of the initial fruit fly population, x k is the k-th fruit fly individual, K is the population size, and Pos k represents the distribution scheme of the e-commerce commodity data of the k-th individual in each storage area, and Node k represents the initial configuration scheme of the storage node resources corresponding to the k-th individual;

[0130] S33. Combine the adaptive olfactory guidance strategy of the e-commerce commodity data storage access priority, and dynamically update the fruit fly individual position based on the e-commerce commodity data storage optimization objective function G opt as follows:

[0131]

[0132] Among them, is the update scheme of the position of the k-th fruit fly individual in the next iteration, μ is the global optimization sensitivity coefficient, γ is the access priority sensitivity factor, is the access priority of the j-th storage area defined in claim 3, is the optimal storage position scheme in the current iteration, η is the random perturbation adjustment coefficient, is the Lévy flight random step size combined with the adjustment of the access priority of the storage area Z j , $f_{k}$ is the optimized objective function value of e-commerce commodity data storage corresponding to the $k$-th individual. $f_{best}$ is the optimized objective function value of e-commerce commodity data storage for the current population.

[0133] S34. Construct a memory-enhanced local aggregation optimization strategy based on the historical access hotspots information of the e-commerce platform, and further update the positions of the fruit fly individuals:

[0134]

[0135] where $\overrightarrow{X}_{k}^{m+1}$ is the e-commerce commodity data storage location scheme after memory enhancement, $\theta$ is the historical hotspot sensitivity coefficient, $h(Z_{k})$ represents the historical access frequency of the storage area $Z_{k}$ corresponding to the $k$-th individual, j and $h(Z_{j})$ is the historical average access frequency of the $j$-th storage area. $\overrightarrow{X}_{local}^{*}$ is the locally optimal storage location determined based on the historical hotspot e-commerce commodity data access frequency within the storage area $Z_{k}$. j

[0136] S35. Repeat the above steps S33 and S34. Dynamically adjust the global optimization sensitivity coefficient $\mu$, the random perturbation adjustment coefficient $\eta$ and the historical hotspot sensitivity coefficient $\theta$ according to the real-time fitness change of the e-commerce commodity data storage structure, and continuously iteratively update the fruit fly population individual position scheme until the e-commerce commodity data storage optimization objective function $G$ opt converges, and output the finally optimized e-commerce commodity data storage scheme:

[0137] $X$ final $= {x$ best $| x$ best $= (Pos$ best , Node$ best ) \}$;

[0138] where $X$ final represents the finally optimized fruit fly population individual position scheme, $Pos$ best is the storage distribution position of the optimized e-commerce commodity data in each storage area, and $Node$ best is the optimized storage node resource allocation scheme.

[0139] This embodiment uses an improved fruit fly optimization algorithm to optimize the storage of e-commerce product data, achieving dynamic adaptive adjustment in terms of storage node resource allocation and data distribution optimization. Through an optimization strategy that combines global search and local tuning, it ensures that the data storage structure conforms to the optimal storage layout, improving the query response speed. An access priority adaptive olfactory guidance strategy is introduced, enabling the data storage optimization to effectively adapt to the characteristics of e-commerce data access, enhancing the storage performance, data access efficiency, and storage resource utilization rate of the system, and reducing data access latency.

[0140] In this embodiment, S4 includes the following steps:

[0141] S41. Based on the standardized e-commerce product data set D std Extract the high-dimensional feature information of e-commerce product data and construct a high-dimensional feature matrix suitable for indexing e-commerce product data:

[0142]

[0143] Among them, V high represents the high-dimensional feature matrix of e-commerce product data, v i is the high-dimensional feature vector of the i-th e-commerce product data, v id represents the d-th feature value of the i-th e-commerce product data, D is the total dimension of the high-dimensional features, and N is the total amount of e-commerce product data;

[0144] S42. According to the historical access frequency characteristics of e-commerce platform product data, introduce an e-commerce product data access heat adaptive hash weight adjustment strategy, construct a method for generating a locality-sensitive hash index vector for e-commerce product data with adaptive access heat, and define the hash index vector of the i-th e-commerce product data as:

[0145]

[0146] Among them, H i is the low-dimensional hash index vector of the i-th e-commerce product data, h il is the l-th hash index value of the i-th e-commerce product data, L is the dimension of the locality-sensitive hash index vector of e-commerce product data, W l is the random projection initial weight vector of the l-th hash function, is the l-th hot spot adjustment weight vector adaptively learned based on the historical access hot spot characteristics of e-commerce product data, F i is the access frequency of the i-th e-commerce product data, ρ is the adjustment coefficient of the e-commerce platform access heat on the hash weight adjustment, b l is the offset value of the l-th hash function;

[0147] S43. Calculate the characteristic mean and covariance matrix of the data in each storage area Z based on the set of e-commerce product data storage areas Z j :

[0148]

[0149] Among them, μ Zj represents the average characteristic vector of the e-commerce product data in the storage area Z j , and is the covariance matrix of the e-commerce product data in the storage area Z j , is the number of e-commerce product data in the storage area Z j , and T is the transpose;

[0150] And optimize the approximate retrieval accuracy of e-commerce product data, and define an adaptive hash distance that combines the differences in e-commerce product data types and the access priorities of storage areas:

[0151]

[0152] Among them, D(H i , H j ) is the hash distance between the i-th and j-th e-commerce product data, and h il , h jl are the l-th hash index values corresponding to the e-commerce product data i and j respectively, is the weight adjustment coefficient of the l-th hash index bit adaptively determined by combining the differences in e-commerce product data types and the access priority of the j-th storage area Z j ;

[0153]

[0154] Among them, β is the adjustment weight of the storage area characteristics for hash distance calculation, and are the mean and variance of the l-th dimension feature in the storage area Z j respectively;

[0155] S44. Based on the set of hash index vectors and the adaptive hash distance constructed in steps S42 and S43, construct a multi-level locality-sensitive hash index table suitable for e-commerce product data location and approximate retrieval:

[0156]

[0157] Among them, is the multi-level locality-sensitive hash index table of the e-commerce product data in the j-th storage area, and H iis the hash index vector of the i-th e-commerce product data, Addr i is the storage address of the e-commerce product data in storage area Z j and Layer i is the hash index level divided in the storage area based on the access frequency and data type of the e-commerce product data.

[0158] In this embodiment, the local sensitive hashing technology is used to optimize the index of e-commerce product data, realizing the low-dimensional and efficient mapping of high-dimensional product data, improving the data retrieval speed and accuracy. The access heat adaptive hashing weight adjustment strategy is adopted to enable the index update to adapt to the access characteristics of e-commerce product data, reducing the query overhead of high-frequency data. The multi-level local sensitive hashing index table structure is introduced to hierarchize the data indexes with different access priorities, improving the accuracy and calculation efficiency of index matching, and ensuring the efficiency of the data storage system in high-concurrency query scenarios.

[0159] In this embodiment, S5 includes the following steps:

[0160] S51. Based on the e-commerce product data storage scheme Pos best and the multi-level local sensitive hashing index table establish the global mapping relationship of the e-commerce information storage management system, and define the global data mapping relationship set:

[0161] M store = {(v i , Pos best,i , H i , Addr i , Layer i ) | v i ∈ D std};

[0162] where M store is the global data mapping relationship set in the e-commerce information storage management system, H i is the hash index vector corresponding to the e-commerce product data, Addr i is the physical storage address of the e-commerce product data, and Layer i is the level information of the e-commerce product data in the multi-level local sensitive hashing index table;

[0163] S52. According to the dynamic access frequency of the e-commerce product data, dynamically adjust the storage location of the e-commerce product data, and define the storage migration optimization rule:

[0164]

[0165] where Pos new,iFor the adjusted storage location of e-commerce commodity data, α new For the storage migration weight parameter, P access,i For the real-time access priority of the i-th e-commerce commodity data For the storage area Z j The average value of the historical access priority

[0166] S53. Adjust the hash index table according to the change of the access mode of e-commerce commodity data Define the dynamic update rule of the hash index:

[0167]

[0168] Among them, For the updated hash index table at the (t + 1)-th moment Indicates the newly added index items due to new data or data storage migration Is the index deletion item due to data migration or deletion

[0169] S54. Construct an access path optimization model for e-commerce commodity data in combination with the global storage management system of e-commerce commodity data. The access path optimization model of e-commerce commodity data considers the data storage location Pos new,i And the hash index table after index update Define the e-commerce commodity data access optimization objective function:

[0170]

[0171] Among them, G access For the e-commerce commodity data access optimization objective function, T query,i For the query response time of the i-th e-commerce commodity data, C query,i For the computational complexity of the i-th e-commerce commodity data query, D index,i For the influence degree of index update on the query path:

[0172]

[0173] Among them, γ1 is the influence weight coefficient of index adjustment on the access optimization objective function;

[0174] The influence degree D of index update on the query path index,i Reflects the influence of the change of index update on the data access path. The larger its value, the more intense the index update, which will cause changes in the data query path;

[0175] S55. According to the optimized data storage scheme Pos new,i And the e-commerce commodity data access optimization objective function Gaccess Build the final e-commerce information storage management system and perform iterative optimization of the storage structure:

[0176]

[0177] Among them, is the final storage location of the i-th e-commerce product data after the (t + 1)-th iteration, λ1 is the global storage adjustment step coefficient, is the access optimization target value of the i-th e-commerce product data at the t-th iteration, and are the minimum and maximum values of the e-commerce product data access optimization objective function in the t-th iteration respectively, Pos new,i is the latest storage location, is the final storage location of the previous round of optimization, ξ is the influence factor of index adjustment on the update of the storage scheme. If the influence degree of index update on the query path is greater than the preset value, then adjust the storage location to adapt to the change of the index and optimize the retrieval efficiency.

[0178] This embodiment constructs a unified e-commerce information storage management system, enables the collaborative optimization of the storage optimization scheme and the data index structure, and improves the overall performance of the system. By using the data access optimization objective function, combined with the dynamic adjustment of the storage location and the optimization of index update, the data query speed and system response efficiency of the e-commerce platform are improved. By dynamically adjusting the storage location through the index update influence factor and optimizing the query path, the retrieval accuracy is improved, ensuring that the data storage and index system always maintains the optimal state in the changing data environment and realizing the efficiency and stability of e-commerce data management.

[0179] An e-commerce information storage system based on big data analysis, used to execute an e-commerce information storage method based on big data analysis, includes the following modules:

[0180] A storage data preprocessing module, used to receive e-commerce product data in the e-commerce platform and perform standardization processing on the data to generate a standardized e-commerce product data set;

[0181] A hierarchical storage management module, used to construct a hierarchical storage structure based on the standardized e-commerce product data set, divide the e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of the e-commerce product data, and set the access priorities of each storage area;

[0182] An improved fruit fly optimization storage optimization module is used to optimize the distribution of e-commerce commodity data in each storage area and the storage node resource configuration based on the improved fruit fly optimization algorithm, and establish an optimized e-commerce commodity data storage scheme;

[0183] An index construction and optimization module is used to construct an e-commerce commodity data index for the standardized e-commerce commodity data set by using the locality-sensitive hashing technique, generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce commodity data, and establish a multi-level locality-sensitive hash index table for the e-commerce commodity data;

[0184] An information storage management module is used to construct a unified e-commerce information storage management system by combining the optimized e-commerce commodity data storage scheme and the multi-level locality-sensitive hash index table of the e-commerce commodity data.

[0185] Example 1:

[0186] On January 10, 2024, during the New Year's Goods Festival promotion of a large e-commerce platform, the technical team noticed that the query response time of commodity searches increased significantly during peak hours. The response time of some query requests even exceeded 2 seconds, resulting in a decline in the user search experience, slow page loading, and a sharp increase in complaints. Through log analysis, the system operation and maintenance team found that the database access bottleneck mainly occurred in the query module of the commodity details page. The queries in this module involved a large amount of high-dimensional data, including basic commodity information, user historical browsing records, and commodity ratings. Moreover, the utilization rate of database storage resources was uneven, and some frequently accessed data was stored in inefficient storage areas, affecting the query efficiency.

[0187] To verify the feasibility of the method of the present invention, the technical team decided to use the method of the present invention to optimize the storage and accelerate the query of e-commerce commodity data. The optimization plan was divided into two stages: the first stage used the improved fruit fly optimization algorithm for storage optimization, and the second stage used the locality-sensitive hash index to optimize the query structure.

[0188] At 1:00 am on January 11, 2024, in order to minimize the impact on user queries, the technical team performed data storage optimization on the database server. First, statistical analysis was carried out on the historical transaction data from January 2023 to January 2024, and it was found that 80% of the query requests were concentrated on the commodity data in the most recent three months, while more than 60% of the historical data (such as transaction records two years ago) was hardly accessed, but this data was still stored on high-performance storage media, occupying a large amount of resources.

[0189] Statistical analysis was carried out on the commodity access frequency from December 2023 to January 2024, and the data was divided into high-frequency access data (the top 15%), medium-frequency access data (30%), and low-frequency access data (55%) according to the access frequency.

[0190] High-frequency access data is stored in the NVMe SSD storage area, medium-frequency data is stored in the SATA SSD storage area, and low-frequency data is stored in the HDD storage area.

[0191] Adopt a dynamic adjustment strategy using the fruit fly optimization algorithm, so that when the access volume of a certain commodity increases by more than 50% within 24 hours, the data storage location of this commodity will be automatically adjusted to a higher-priority storage area.

[0192] On January 12, 2024, after storage optimization, the query response time of the product details page decreased significantly.

[0193] Taking the "New Year's Special Blessing Bag" as an example, the average query response time of this commodity on January 10, 2024 was 1.8 seconds. When querying the same commodity on January 12, 2024 after optimization, the response time was shortened to 0.9 seconds, and the optimization amplitude reached 50%.

[0194] The average response time during low-frequency data access decreased from 1.5 seconds to 1.1 seconds, and the overall data access efficiency improved.

[0195] On January 15, 2024, the technical team further optimized the index structure of commodity queries to solve the problem of slow retrieval of high-dimensional feature data. For example, during the "Spring Festival Promotion", when users search for "New Year's Goods Gift Boxes", the platform needs to match users' purchase preferences, product ratings, and historical transaction record information. The original index query method took a long time and affected the user experience.

[0196] At 22:00 on the evening of January 15, 2024, the technical team began to reconstruct the index of the commodity database.

[0197] Use the locality-sensitive hashing technique to build commodity indexes, and adjust the hashing index weights in combination with the commodity access heat, so that the index priority of frequently queried commodities is improved, and the query calculation amount is reduced.

[0198] For example, for commodities that have been queried more than 10,000 times in the past three days, their index weights will be automatically increased, so that the query preferentially matches in the front-level layer of the hashing index table, reducing the query time.

[0199] On January 16, 2024, when querying the "Spring Festival Big Gift Package", the system response time decreased from the original 700ms to 410ms, and the query efficiency increased by 41.4%.

[0200] Taking the query log data at 20:00 on the evening of January 16, 2024 as an example:

[0201] The query time of the traditional MongoDB B+ tree index: average 920ms.

[0202] Query time after optimizing with the LSH index of the present invention: average 520 ms.

[0203] When searching for "Spring Festival gift box", the query success rate for this keyword on January 14, 2024 was 92.4% (partial queries timed out). After optimization, the query success rate increased to 99.1% on January 16, 2024, and the query accuracy improved.

[0204] On January 18, 2024, the technical team of the e-commerce platform monitored in the background that the query volume of a popular product, "Spring Festival gift box model A", suddenly increased from 3:00 am to 5:00 am, with the cumulative number of queries reaching 12,000 times, an 85% increase compared to 24 hours ago. The fruit fly optimization algorithm automatically identified the sharp increase in the query popularity of this product at 4:30 am and triggered the storage migration mechanism to migrate its data storage from SATA SSD to NVMe SSD, improving the query priority.

[0205] At 12:00 noon on January 18, 2024, the operation team found that the click-through rate of this product on the home page recommendation position increased from the original 2.8% to 5.3%, and the average user stay time increased by 30%, directly driving the sales growth.

[0206] Through the experiment from January 10, 2024 to January 20, 2024, the method of the present invention verified its effectiveness on an actual e-commerce platform and conducted a comparative analysis with traditional methods. The experimental data is as follows:

[0207] Evaluation metrics Traditional method Method of the present invention Improvement rate Average query response time 920ms 520ms 43% Storage resource utilization rate 70% 88% 25% System stability during peak periods Prone to overload Steady - Data storage migration overhead High Low 35%

[0208] The method of the present invention optimizes the data storage structure by improving the fruit fly optimization algorithm, enabling high-frequency data to be stored on efficient storage media and improving the query response speed; combines the locality-sensitive hashing index to optimize the query structure, making the retrieval of high-dimensional data more efficient; and automatically adjusts the data storage location through a dynamic storage migration strategy to improve the storage resource utilization rate. The experimental results show that the method of the present invention has significantly improved in terms of query response speed, storage resource utilization rate, and system stability, and has high practical value and market application prospects.

[0209] The present invention uses an improved fruit fly optimization algorithm to globally optimize and locally tune the storage distribution of e-commerce product data. By combining the access priority of e-commerce product data, the adaptive olfactory guidance strategy, and historical access hot spot information, it realizes the optimal distribution of e-commerce product data in the storage area, and introduces an access frequency sensitive factor and a historical hot spot enhancement strategy, enabling the fruit fly optimization algorithm to dynamically adapt to different access patterns of e-commerce product data during the data distribution adjustment process, thereby reducing the load imbalance of storage nodes and improving the concurrent response ability of data access.

[0210] The present invention introduces the locality-sensitive hashing technology and combines it with an adaptive weight adjustment strategy for the access popularity of e-commerce product data to construct a multi-level hash index table. By adaptively adjusting the weights of the hash index, hot data can be preferentially mapped to high-priority storage areas, thereby improving the query efficiency.

[0211] The present invention constructs an adaptive hierarchical storage structure based on a standardized e-commerce product data set and adjusts the data storage location according to the real-time access pattern of e-commerce data through a dynamic storage migration optimization strategy, so as to achieve a reasonable allocation of storage resources. Specifically, based on the storage migration optimization rules based on access priorities, e-commerce product data can be dynamically adjusted between hot and cold data storage areas, avoiding the storage of hot data in low-efficiency storage areas and improving the access performance.

[0212] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An e-commerce information storage method and system based on big data analysis, characterized in that, Including the following steps: S1. Format the e-commerce product data collected from the e-commerce platform to form a standardized e-commerce product data set; S2. Based on the standardized e-commerce product data set, construct a hierarchical storage structure, divide the e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of the e-commerce product data, and preset the access priorities of each storage area; S3. Use the improved fruit fly optimization algorithm to perform global search and local tuning on the hierarchical storage structure, optimize the distribution of e-commerce product data in each storage area and the resource configuration of storage nodes, and form an optimized e-commerce product data storage solution; S4. Use the locality-sensitive hashing technique to construct an e-commerce product data index for the standardized e-commerce product data set, generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data, and establish a unified multi-level locality-sensitive hashing index table for e-commerce product data in each storage area; S5. Combine the optimized e-commerce product data storage solution with the multi-level locality-sensitive hashing index table to form a unified e-commerce information storage management system.

2. The e-commerce information storage method and system based on big data analysis according to claim 1, characterized in that, The S1 includes the following steps: S11. Collect e-commerce product data from the e-commerce platform. The e-commerce product data includes product basic information, classification information, historical transaction records and user interaction information, and construct an e-commerce product data set; S12. Format the e-commerce product data set so that all data items in the e-commerce product data set meet the unified data storage format standard; S13. Perform consistency check on the formatted e-commerce product data set so that all data items in the e-commerce product data set meet the requirements of integrity, correctness and uniqueness. Define the formatted standardized e-commerce product data set as: ; Among them, is the standardized e-commerce product dataset, is the th formatted product data record, , , , respectively represent the basic product information, product classification information, historical transaction records, and user interaction information after standardization, is the total amount of product data.

3. The method and system for storing e-commerce information based on big data analysis according to claim 2, wherein, The S2 includes the following steps: S21. Based on the standardized e-commerce product dataset , count the access frequency, data type, and storage resource requirements of e-commerce product data, and establish an e-commerce product data feature matrix: ; Among them, is the e-commerce product data feature matrix, represents the feature set of the th e-commerce product data record, is the access frequency of the th e-commerce product data, is the data type of the th e-commerce product data, is the storage resource requirement of the th e-commerce product data; S22. Determine the storage area division rules for e-commerce product data according to the e-commerce product data feature matrix and map the e-commerce product data to different storage areas: ; Among them, represents the set of e-commerce commodity data storage area, is the th storage area, represents the total number of storage areas, , , respectively represent the access frequency range, data type range and storage resource allocation rule corresponding to the th storage area; S23. For each storage area Set a preset access priority, and define the access priorities of each storage area as follows: ; Among them, represents the access priority of the th storage area, , , are respectively the normalized average values of the access frequency, data type, and storage resource requirements of e-commerce commodity data in the th storage area, , , are respectively the weight coefficients of the access frequency, data type, and storage resource requirements; S24. According to the storage area access priority Sort the set of storage areas to complete the construction of the hierarchical storage structure of e-commerce commodity data.

4. An e-commerce information storage method and system based on big data analysis according to claim 3, characterized in that, The S3 includes the following steps: S31. According to the e-commerce product data hierarchical storage structure, with the goal of reducing the response latency of the e-commerce platform, reducing data redundancy and improving the load balance of storage nodes, establish an e-commerce product data storage optimization objective function: ; Among them, is the optimization objective function for e-commerce commodity data storage, is the th storage area average access latency of e-commerce commodity data in, is the maximum value of the access latency of e-commerce commodity data in each storage area, is the th redundancy data ratio of e-commerce commodity data in the storage area, is the maximum value of the redundancy ratio of e-commerce commodity data in each storage area, is the th imbalance degree of the storage node load in the storage area, is the maximum value of the imbalance degree of the storage node load in each storage area, , , are the optimization weight coefficients of access latency, data redundancy, and load balancing in the e-commerce platform, respectively; S32. Initialize the population of the improved fruit fly optimization algorithm applicable to the optimization of the e-commerce product data storage structure, and establish the definition of the initial fruit fly individual position as: ; Among them, is the set of individual positions of the initial Drosophila population, is the th Drosophila individual, is the population size, represents the distribution scheme of the e-commerce commodity data of the th individual in each storage area, represents the initial configuration scheme of the storage node resources corresponding to the th individual; S33. Adaptive olfactory guidance strategy combined with the storage access priority of e-commerce product data, with the optimization objective function of e-commerce product data storage Based on this, dynamically update the positions of fruit fly individuals: ; Among them, is the update scheme of the position of the th Drosophila individual in the next iteration, is the global optimization sensitivity coefficient, is the access priority sensitivity factor, is the access priority of the th storage area defined in claim 3, is the optimal storage location scheme in the current iteration, is the random perturbation adjustment coefficient, is the Levy flight random step size adjusted by combining the access priority of the storage area , is the value of the optimization objective function for the storage of e-commerce commodity data corresponding to the th individual, is the optimal value of the optimization objective function for the storage of e-commerce commodity data in the current population; S34. According to the historical access hot spot information of the e-commerce platform, construct a memory-enhanced local aggregation optimization strategy to further update the fruit fly individual position: ; Among them, is the storage location scheme of e-commerce commodity data after memory enhancement, is the historical hot spot sensitivity coefficient, represents the historical access frequency of the storage area corresponding to the th individual, is the historical average access frequency of the th storage area, is the local optimal storage location determined based on the historical access frequency of hot spot e-commerce commodity data in the storage area ; Repeat the above steps S33 and S34, and dynamically adjust the global optimization sensitivity coefficient according to the real-time fitness change of the e-commerce commodity data storage structure , the random perturbation adjustment coefficient and the historical hotspot sensitivity coefficient , and continuously iteratively update the individual position scheme of the fruit fly population until the e-commerce commodity data storage optimization objective function converges, and output the finally optimized e-commerce commodity data storage scheme: ; Among them, represents the individual position solution of the finally optimized Drosophila population, is the storage distribution position of the optimized e-commerce commodity data in each storage area, is the optimized storage node resource allocation solution.

5. A method and system for storing e-commerce information based on big data analysis according to claim 4, characterized in that, The S4 includes the following steps: S41. Based on the standardized e-commerce commodity dataset Extract the high-dimensional feature information of e-commerce commodity data and construct a high-dimensional feature matrix suitable for indexing e-commerce commodity data: ; in, represents the high-dimensional feature matrix of e-commerce commodity data, For the high-dimensional feature vectors of e-commerce product data, Indicates E-commerce product data dimensional eigenvalue, is the total dimension of the high-dimensional features, is the total amount of e-commerce commodity data; S42. According to the historical access frequency characteristics of e-commerce platform commodity data, introduce an adaptive hash weight adjustment strategy for the access heat of e-commerce commodity data, construct a method for generating a locality-sensitive hash index vector for e-commerce commodity data with adaptive access heat, and define the hash index vector of the th e-commerce commodity data as follows: ; Among them, is the low-dimensional hash index vector of the th e-commerce product data, is the th hash index value of the th e-commerce product data. L is the dimension of the locality-sensitive hash index vector of the e-commerce product data, is the random projection initial weight vector of the th hash function, is the th hot-spot adjustment weight vector adaptively learned based on the historical access hot-spot features of the e-commerce product data, is the access frequency of the th e-commerce product data, is the adjustment coefficient of the hash weight adjustment by the access heat of the e-commerce platform, is the th offset value of the hash function; S43. Based on the set of e-commerce commodity data storage areas , calculate the characteristic mean and covariance matrix of the data in each storage area : ; ; Among them, represents the average eigenvector of e-commerce product data in the storage area, is the covariance matrix of e-commerce product data in the storage area, is the number of e-commerce product data in the storage area, and T is the transpose; And optimize the approximate retrieval accuracy of e-commerce product data, and define an adaptive hash distance that combines the e-commerce product data type difference and the storage area access priority: ; Among them, is the th hash distance from the th e-commerce product data, and are the th hash index values corresponding to the e-commerce product data is the th hash index bit weight adjustment coefficient adaptively determined by combining the differences in e-commerce product data types and the th access priority of the storage area: ; Among them, is the adjusted weight for the hash distance calculation of the storage area feature, and are respectively the mean and variance of the -dimensional feature within the storage area; S44. Based on the hash index vector set and the adaptive hash distance constructed in steps S42 and S43, construct a multi-level locality-sensitive hashing index table applicable to the positioning and approximate retrieval of e-commerce product data: ; Among them, is the multi-level locality-sensitive hashing index table of e-commerce product data for the th storage area, is the hashing index vector of the th e-commerce product data, is the storage address of the e-commerce product data in the storage area , is the hashing index level divided within the storage area based on the access frequency and data type of the e-commerce product data.

6. The method and system for storing e-commerce information based on big data analysis according to claim 5, characterized in that, The S5 includes the following steps: S51. E-commerce commodity data storage solution and multi-level locality-sensitive hashing index table , establish the global mapping relationship of the e-commerce information storage management system, and define the global data mapping relationship set: ; Among them, is the global data mapping relationship set in the e-commerce information storage management system, is the hash index vector corresponding to the e-commerce product data, is the physical storage address of the e-commerce product data, is the hierarchical information of the e-commerce product data in the multi-level locality-sensitive hashing index table; S52. Dynamically adjust the storage location of e-commerce product data according to the dynamic access frequency of e-commerce product data, and define storage migration optimization rules: ; Among them, is the storage location of the adjusted e-commerce product data, is the storage migration weight parameter, is the real-time access priority of the th e-commerce product data, is the average value of the historical access priorities of the storage area; S53. Adjust the hash index table according to the change of the access mode of e-commerce commodity data , define the dynamic update rule of the hash index: ; Among them, is the updated hash index table at the th moment, indicating the newly added index items due to newly added data or data storage migration, is the index deletion item due to data migration or deletion; S54. Construct an access path optimization model for e-commerce product data by combining the global storage management system of e-commerce product data. The access path optimization model for e-commerce product data takes into account the data storage location and the hash index table after index update to define the access optimization objective function for e-commerce product data: ; Among them, is the optimization objective function for e-commerce product data access, is the query response time of the th e-commerce product data, is the th computational complexity of the e-commerce product data query, is the degree of influence of index update on the query path: ; Among them, is the influence weight coefficient of the index adjustment on the access optimization objective function; Degree of influence of index update on query path Reflects the impact of changes in index updates on the data access path. The larger the value, the more drastic the index update, which will cause changes in the data query path; S55. According to the optimized data storage solution and the optimization objective function of e-commerce commodity data access Construct the final e-commerce information storage management system and perform the optimization iteration of the storage structure: ; Among them, is the final storage location of the th e-commerce product data after the th iteration, is the global storage adjustment step coefficient, is the th access optimization target value of the e-commerce product data at the th iteration, and are respectively the minimum and maximum values of the access optimization objective function of the e-commerce product data in the th iteration, is the latest storage location, is the final storage location of the previous round of optimization, is the influence factor of the index adjustment on the update of the storage scheme. If the influence degree of the index update on the query path is greater than the preset value, the storage location is adjusted to adapt to the change of the index and optimize the retrieval efficiency.

7. An e-commerce information storage system based on big data analysis, which is used to execute an e-commerce information storage method based on big data analysis described in any one of claims 1-6, characterized in that, Including the following modules: A storage data preprocessing module, which is used to receive e-commerce product data in an e-commerce platform, perform standardization processing on the data, and generate a standardized e-commerce product data set; A hierarchical storage management module, which is used to construct a hierarchical storage structure based on the standardized e-commerce product data set, divide the e-commerce product data into different storage areas according to the access frequency, data type and storage resource allocation requirements of the e-commerce product data, and set the access priorities of each storage area; An improved fruit fly optimization storage optimization module, which is used to optimize the distribution of e-commerce product data in each storage area and the storage node resource configuration based on the improved fruit fly optimization algorithm, and establish an optimized e-commerce product data storage scheme; An index construction and optimization module, which is used to construct an e-commerce product data index for the standardized e-commerce product data set by using the locality-sensitive hashing technique, generate a low-dimensional hash index based on the high-dimensional feature information of the e-commerce product data, and establish a multi-level locality-sensitive hashing index table for the e-commerce product data; An information storage management module, which is used to construct a unified e-commerce information storage management system by combining the optimized e-commerce product data storage scheme with the multi-level locality-sensitive hashing index table of the e-commerce product data.

Citation Information

Patent Citations

  • Edge data storage method based on aging intimacy model

    CN115865712A

  • Reservoir area layout method and device based on category association rule and storage medium

    CN117333113A

  • Image similarity search via hashes with expanded dimensionality and sparsification

    US20190171665A1

Cited By

  • Preprocessed data partition storage method and system based on big data analysis

    CN121764414A