Private domain e-commerce data search method and system based on big data
By analyzing the user's search feature information, semantic extraction and expansion, combining segmented product quantitative compression and double-layer compact indexing, the problem that the existing private domain e-commerce platform search methods cannot capture the potential intentions of users, and achieve accurate search results matching and improve user experience.
Patent Information
- Application Number
- CN202510549594.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The search method of existing private domain e-commerce platforms is based only on literal matching, and cannot capture the potential intentions and associated information of users, resulting in users requiring repeated searches for multiple times.
By obtaining user search feature information, including behavioral characteristics and purchase characteristics, analyzing user's style preferences and purchase preferences, semantic extraction and expansion, combining segmented product quantization compression and double-layer compact indexing, accurate matching of user potential intentions is achieved.
It realizes accurate capture of users' potential intentions, reduces the number of repeated searches by users, and improves the relevance and user experience of search results.
Smart Images

Figure CN120067461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of private domain e-commerce, and in particular to a method and system for searching private domain e-commerce data based on big data. Background Art
[0002] Private domain e-commerce refers to a business model that relies on the traffic pools owned by individuals or enterprises (such as WeChat official accounts, mini-programs, communities, APPs, etc.) to carry out commodity sales and service provision. Different from traditional e-commerce platforms (such as Taobao and JD.com), private domain e-commerce emphasizes more on establishing and maintaining direct connections with consumers to form a stable customer group, so as to maximize long-term value. Existing private domain e-commerce platforms usually rely on basic keywords for searching, and this search method only returns results based on literal matching, ignoring the potential intentions of users (for example, when a user searches for "summer dresses", the traditional method only returns results based on literal matching, ignoring potential user intentions such as "breathable fabric" and "vacation style"), and cannot give accurate associated information (such as breathable fabric), resulting in users needing to repeat searches multiple times. There is a need for a method for searching private domain e-commerce data based on big data to solve the above problems. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for searching private domain e-commerce data based on big data to solve the technical problems raised in the above background art.
[0004] To achieve the above purpose, the present invention provides the following technical solutions: A method for searching private domain e-commerce data based on big data, comprising: Obtaining user search feature information of a private domain e-commerce platform, where the user search feature information includes behavior feature information and purchased commodity feature information; Obtaining browsing feature information and search keywords according to the behavior feature information, and obtaining multiple browsing commodity click times and multiple browsing click stay times according to the browsing feature information, and obtaining style preference features according to the multiple browsing commodity click times and the multiple browsing click stay times; Performing semantic extraction on the search keywords to obtain explicit semantic features, and obtaining extended semantic features according to the explicit semantic features and style preference features; Obtaining repurchase data information and timeliness information according to the purchased commodity feature information, and obtaining purchase preference features according to the repurchase data information and the timeliness information; Performing segmented product quantization compression on the extended semantic features and the purchase preference features to obtain compressed semantic association features, and performing simultaneous double-layer compact indexing on the explicit semantic features and the compressed semantic association features to obtain explicit results and association results; Visualize the private domain e-commerce search results according to the explicit results and the associated results.
[0005] Preferably, the step of obtaining the style preference features according to the multiple click times of the browsed products and the multiple click-and-stay times of the browsed clicks includes: Sort the multiple click times of the browsed products in descending order to obtain a sorted list of click times of the browsed products; Sort the multiple click-and-stay times of the browsed clicks in descending order to obtain a sorted list of click-and-stay times of the browsed clicks; Obtain the first product corresponding to the click time of the first browsed product according to the sorted list of click times of the browsed products; Obtain the second product corresponding to the first click-and-stay time of the browsed clicks according to the sorted list of click-and-stay times of the browsed clicks, and perform dynamic load factor constraint on the second product and the first product to obtain an equilibrium clustering cluster set, and perform centroid screening on the equilibrium clustering cluster set based on the K-means clustering algorithm to obtain product centroids; Perform fine-grained semantic feature screening on the product centroids based on the attention mechanism to obtain multiple fine-grained semantic features; Extract the first keywords from the multiple fine-grained semantic features based on the BERT language model, and capture the context association of the first keywords based on the multi-head self-attention mechanism to obtain keyword semantic texts; Extract the style preference features of the keyword semantic texts based on a preset text library.
[0006] Preferably, the step of performing semantic extraction on the search keywords to obtain explicit semantic features and obtaining extended semantic features according to the explicit semantic features and style preference features includes: Perform word segmentation on the search keywords to obtain multiple independent words of the search keywords, and convert the multiple independent words into word vectors based on the Word2Vec model to obtain an independent word set vector, and define the semantic features corresponding to the independent word set vector as explicit semantic features; Map the style preference features to a style semantic vector aligned with the dimension of the independent word set vector through a fully connected layer; Perform comparison and screening on the independent word set vector and the style semantic vector based on the cosine similarity model to obtain an independent word-style semantic subset vector; Perform vector splicing on the independent word-style semantic subset vector and the style semantic vector based on a preset time series to obtain a spliced vector; Input the spliced vector into a preset bidirectional LSTM network for context semantic enhancement to obtain an enhanced spliced vector, and use the semantic features of the enhanced spliced vector as extended semantic features.
[0007] Preferably, the step of obtaining repurchase data information and timeliness information according to the purchased commodity feature information, and obtaining purchase preference features according to the repurchase data information and timeliness information includes: Obtain the historical purchase timestamp and multiple commodity categories according to the purchased commodity feature information, and obtain the time interval between two adjacent purchases under each same commodity category; Mark the behaviors with multiple time intervals less than a preset threshold as repurchase events, and define the data information corresponding to the repurchase events as repurchase data information; Obtain a time sliding window according to the historical purchase timestamp, obtain a seasonal distribution interval according to the time sliding window, divide the multiple commodity categories into the seasonal distribution interval to obtain seasonally interval-normalized commodity categories, and use the time information of the seasonally interval-normalized commodity categories in the seasonal distribution interval as timeliness information; Perform timeliness screening on the repurchase data information according to the timeliness information to obtain screened repurchase data information, and perform encoding processing on the screened repurchase data information to obtain a screened repurchase data code; Based on the Word2Vec model, convert the screened repurchase data code into a word vector to obtain a screened repurchase data code vector, and use the screened repurchase data code vector as the purchase preference feature.
[0008] Preferably, the step of performing segmented product quantization compression on the extended semantic feature and the purchase preference feature to obtain a compressed semantic association feature includes: Map the extended semantic feature to an extended semantic feature vector; Map the purchase preference feature to a purchase preference feature vector, and splice the extended semantic feature vector and the purchase preference feature vector to obtain a joint vector; Divide the joint vector into multiple sub-vector segments according to a preset length, and perform clustering on each sub-vector segment based on K-means clustering to obtain multiple clustering centroids, where each clustering centroid represents an index; Map the multiple clustering centroids to multiple clustering centroid vectors, perform product quantization on the multiple clustering centroid vectors to obtain a mass point vector value, and use the semantics corresponding to the mass point vector value as the compressed semantic association feature.
[0009] Preferably, the step of performing simultaneous double-layer compact indexing on the explicit semantic feature and the compressed semantic association feature to obtain an explicit result and an association result includes: Construct the explicit semantic features based on the inverted index of explicit features to obtain the first-layer index, and map the first-layer index to hash buckets based on Locality-Sensitive Hashing (LSH) to obtain the first-layer hash index. Use the result corresponding to the first-layer hash index as the explicit result; Construct the second-layer index based on the compressed feature map index for the compressed semantic association features, and construct the second-layer index into a multi-layer graph index based on the Hierarchical Navigable Small World (HNSW) graph algorithm, where each layer of the graph index corresponds to a product identifier; Use the result corresponding to the multi-layer graph index as the association result.
[0010] This application also provides a private domain e-commerce data search system based on big data, including: A first acquisition module for acquiring user search feature information of a private domain e-commerce platform, where the user search feature information includes behavioral feature information and purchased product feature information; A second acquisition module for acquiring browsing feature information and search keywords according to the behavioral feature information, acquiring the click times of multiple browsed products and the dwell time of multiple browsing clicks according to the browsing feature information, and acquiring style preference features according to the click times of multiple browsed products and the dwell time of multiple browsing clicks; An extraction module for performing semantic extraction on the search keywords to obtain explicit semantic features, and obtaining extended semantic features according to the explicit semantic features and style preference features; A third acquisition module for acquiring repurchase data information and timeliness information according to the purchased product feature information, and acquiring purchase preference features according to the repurchase data information and timeliness information; A fourth acquisition module for performing segmented product quantization compression on the extended semantic features and the purchase preference features to obtain compressed semantic association features, and performing simultaneous double-layer compact indexing on the explicit semantic features and the compressed semantic association features to obtain an explicit result and an association result; A presentation module for visually presenting the private domain e-commerce search results according to the explicit result and the association result.
[0011] Preferably, the second acquisition module includes: A first sorting unit for performing a descending sort on the click times of multiple browsed products to obtain a sorted list of click times of browsed products; A second sorting unit for performing a descending sort on the dwell time of multiple browsing clicks to obtain a sorted list of dwell time of browsing clicks; A first acquisition unit for acquiring the first product corresponding to the click times of the top-ranked browsed product according to the sorted list of click times of browsed products; A second acquisition unit, configured to acquire a second product corresponding to the top-ranked browsing click stay time sorting table according to the browsing click stay time sorting table, and perform dynamic load factor constraint on the second product and the first product to obtain an equilibrium clustering cluster set, and perform centroid screening on the equilibrium clustering cluster set based on the K-means clustering algorithm to obtain a product centroid; A screening unit, configured to perform fine-grained semantic feature screening on the product centroid based on an attention mechanism to obtain a plurality of fine-grained semantic features; A first extraction unit, configured to perform keyword extraction on the plurality of fine-grained semantic features based on a BERT language model to obtain a first keyword, and perform context association capture on the first keyword based on a multi-head self-attention mechanism to obtain a keyword semantic text; A second extraction unit, configured to extract a style preference feature of the keyword semantic text based on a preset text library.
[0012] The present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0013] The present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0014] The beneficial effects of the present application are as follows: The present invention first collects user behavior characteristics (search, browsing duration, product click volume) and purchase characteristics (repurchase cycle, timeliness preference), constructs a style preference model and a purchase preference model, quantifies the user's attention to product styles based on browsing data, combines the explicit semantics of search keywords for semantic extension, captures potential demands such as "breathable fabric" and "vacation style", and at the same time uses a segmented product quantization technique to compress the extended semantics and purchase preferences. Through a two-layer compact index, parallel retrieval of explicit demands and associated demands is realized. Finally, the structured search results are visually presented, integrating the basic information of the product and the style extension features to meet the user's potential intention demands, and avoiding the need for the user to repeat searches multiple times. Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present application.
[0016] Figure 2 It is a schematic structural diagram of the system according to an embodiment of the present application.
[0017] Figure 3 It is a schematic internal structure diagram of the computer device according to an embodiment of the present application.
[0018] The realization of the purpose, functional features and advantages of this application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific Embodiments
[0019] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] As Figures 1-3 shown, this application provides a method for searching private domain e-commerce data based on big data, including: S1. Obtain the user search feature information of the private domain e-commerce platform, where the user search feature information includes behavioral feature information and purchased product feature information; S2. Obtain the browsing feature information and search keywords according to the behavioral feature information, obtain the click times of multiple browsed products and the dwell time of multiple browsing clicks according to the browsing feature information, and obtain the style preference feature according to the click times of multiple browsed products and the dwell time of multiple browsing clicks; S3. Perform semantic extraction on the search keywords to obtain the explicit semantic feature, and obtain the extended semantic feature according to the explicit semantic feature and the style preference feature; S4. Obtain the repurchase data information and timeliness information according to the purchased product feature information, and obtain the purchase preference feature according to the repurchase data information and the timeliness information; S5. Perform segmented product quantization compression on the extended semantic feature and the purchase preference feature to obtain the compressed semantic association feature, and perform simultaneous double-layer compact indexing on the explicit semantic feature and the compressed semantic association feature to obtain the explicit result and the association result; S6. Visualize the private domain e-commerce search results according to the explicit result and the association result.
[0021] As described in the above steps S1-S6, since existing private domain e-commerce platforms usually rely on basic keywords for searching, and this search method only returns results based on literal matching, it will ignore the user's potential intentions (for example, when a user searches for "summer dress", the traditional method only returns results based on literal matching, ignoring the user's potential intentions, such as "breathable fabrics" and "vacation style", etc.), and cannot provide accurate related information (breathable fabrics, etc.), which in turn causes the user to repeat the search multiple times. The present invention first obtains the user search feature information of the private domain e-commerce platform, wherein the user search feature information includes behavioral feature information and purchased product feature information. In this way, by collecting the user's behavior and purchase-related information on the private domain e-commerce platform, the foundation is laid for the subsequent in-depth analysis of user needs and preferences. By obtaining these characteristic information, the user's activities on the platform can be fully understood, including search, browsing, purchase and other behaviors, so as to mine the user's potential needs, and then the browsing characteristic information and search keywords are obtained according to the behavioral characteristic information, and multiple browsing product clicks and multiple browsing click dwelling times are obtained according to the browsing characteristic information, and the style preference characteristics are obtained according to the multiple browsing product clicks and multiple browsing click dwelling times, so that the browsing and search related data are extracted from the behavioral characteristic information, and the user's style preference characteristics are further analyzed. This helps to grasp the user's preferences more accurately and provide direction for the expansion and optimization of search results. By analyzing the number of browsing product clicks and dwelling time, the user's attention to different product styles can be understood, so as to infer their style preferences. At the same time, the user's browsing and search behaviors can intuitively reflect their points of interest. By mining these behavioral data, the user's potential style preferences can be discovered, so that the search results are more in line with the user's expectations, and then the search keywords are semantically extracted to obtain explicit semantic features, and extended semantic features are obtained according to the explicit semantic features and style preference features, and the explicit semantic features of the search keywords are mined, and the semantics are extended in combination with the style preference features, so that the search results are more comprehensive and more in line with the user's potential intentions. It can not only match the literal meaning, but also take into account the related semantics that the user may be interested in, thereby improving the accuracy and relevance of the search. Secondly, according to the purchased product feature information, repeat purchase data information and time information are obtained, and according to the repeat purchase data information and time information, purchase preference characteristics are obtained. By analyzing the repeat purchase data and time information in the purchased product feature information, the user's purchase preference characteristics are obtained.This can help the platform understand the purchasing habits and changing needs of users, providing a basis for personalized recommendations and optimizing search results. At the same time, the purchasing behavior is a direct manifestation of user needs. The repurchase data and timeliness information can reflect the degree of user preference for products and the purchase time pattern, which is of great significance for improving the quality of search services. Then, the extended semantic features and the purchase preference features are subjected to segmented product quantization compression to obtain compressed semantic association features, and the explicit semantic features and the compressed semantic association features are simultaneously subjected to a double-layer compact index to obtain explicit results and association results. By segmented product quantization compression, the data dimension is reduced, and the data processing efficiency is improved. At the same time, through double-layer compact indexing, the explicit results and association results can be quickly obtained, making the search results more accurate and efficient, and capable of simultaneously giving users' potential intentions, such as "breathable fabric" and "vacation style". Finally, based on the explicit results and the association results, the search results of the private domain e-commerce are visually presented, and then the explicit results and the extended association results are displayed to users in an intuitive manner, facilitating users to quickly find the required products and enhancing the user experience and shopping efficiency. Visual presentation can enable users to more clearly understand the content and characteristics of the search results, improve users' satisfaction with the platform, and at the same time avoid the problem that users need to repeat searches multiple times.
[0022] In one embodiment, the step S2 of obtaining the style preference features according to the multiple browsing product click times and the multiple browsing click stay times includes: S201. Sort the multiple browsing product click times in descending order to obtain a browsing product click time sorting table; S202. Sort the multiple browsing click stay times in descending order to obtain a browsing click stay time sorting table; S203. Obtain the first product corresponding to the first browsing product click time according to the browsing product click time sorting table; S204. Obtain the second product corresponding to the first browsing click stay time sorting table according to the browsing click stay time sorting table, and perform dynamic load coefficient constraint on the second product and the first product to obtain an equilibrium clustering cluster set, and perform centroid screening on the equilibrium clustering cluster set based on the K-means clustering algorithm to obtain product centroids; S205. Based on the attention mechanism, perform fine-grained semantic feature screening on the product centroids to obtain multiple fine-grained semantic features; S206. Based on the BERT language model, extract the first keywords from the multiple fine-grained semantic features, and capture the context association of the first keywords based on the multi-head self-attention mechanism to obtain keyword semantic texts; S207. Extract the style preference features of the keyword semantic texts based on a preset text library.
[0023] As described in the above steps S201 - S207, the present invention first performs a descending sort on the click times of multiple said browsed products to obtain a sorted list of click times of browsed products. Sorting the click times of the products browsed by the user can intuitively present the differences in the degree of attention of the user to different products. Products with more click times are likely to be of greater interest to the user. Through this sort, these key products can be quickly located, providing a clear direction for subsequent analysis, helping the e-commerce platform accurately grasp the user's interest tendency. Then, a descending sort is performed on the browsing click stay times of multiple said products to obtain a sorted list of browsing click stay times. The browsing click stay time reflects the depth of the user's interest in the product. The longer the stay time, the more interested the user is in the details, introductions, etc. of the product, and the more in-depth understanding of the product. After sorting, the potential interest points of the user in the product can be further explored, assisting in judging the user's true preferences, complementing the sorting results of click times, and more comprehensively depicting the user's interests. Secondly, according to the sorted list of click times of browsed products, the first product corresponding to the click times of the top - ranked browsed product is obtained. In this way, the top - ranked product in the sorted list, that is, the product with the most click times by the user, is determined as one of the products that the user is most concerned about, and it is used as a key sample for analyzing the user's style preference. This product is representative, and its style characteristics are likely to be liked by the user, providing an important clue for subsequent style analysis. Immediately afterwards, according to the sorted list of browsing click stay times, the second product corresponding to the sorted list of browsing click stay times is obtained, and the second product and the first product are subjected to dynamic load coefficient constraint to obtain an equilibrium clustering cluster set. The equilibrium clustering cluster set is subjected to centroid screening based on the K - means clustering algorithm to obtain product centroids. By obtaining the second product with the longest stay time and combining it with the first product, the products corresponding to the user's attention frequency and attention depth are comprehensively considered. The dynamic load coefficient constraint can balance the relationship between the two, avoiding being dominated by a single factor. Among them, the dynamic load coefficient constraint is a dynamic weight allocation mechanism. By weighted adjustment of the contribution weights of different feature dimensions (such as click times, stay time) in the clustering process, the finally generated clustering cluster set can evenly reflect the multi - dimensional user preferences, avoiding a single feature dominating the clustering result. The K - means clustering algorithm performs centroid screening, which can cluster products with similar features into one category and find the centroid representing the features of this category of products, providing a basis for subsequent accurate analysis of the user's style preference and making the analysis results more universal and representative. Then, based on the attention mechanism, fine - grained semantic feature screening is performed on the product centroids to obtain multiple fine - grained semantic features. Among them, the attention mechanism (Attention Mechanism) is a technology in deep learning that simulates the "selective attention" in the human cognitive process. Its core is to focus on the most relevant part of the input information through dynamic weight allocation. In this way, the attention mechanism can focus on the key semantic features of the product centroids and filter out some unimportant information.Multiple sub-semantic features are obtained through screening to more precisely characterize the characteristics of the product centroid, deeply explore specific style elements that users may be interested in, and provide more accurate information for subsequent keyword extraction and style preference determination. For example, in this solution, Dress B ranks first in the browsing click and stay time ranking table and becomes the second product. The dynamic load coefficient constraints are applied to Dress A and Dress B, and then processed through the K-means clustering algorithm. If there are other styles of dresses participating in the clustering, the resulting product centroid may reflect comprehensive features such as floral patterns, body-hugging, and mid-length. These features can better represent Xiao Zhang's overall style preference for dresses. Then, based on the BERT language model, keyword extraction is performed on multiple sub-semantic features to obtain the first keywords, and the context association of the first keywords is captured based on the multi-head self-attention mechanism to obtain keyword semantic text. With the powerful semantic understanding ability of the BERT language model, key information can be accurately extracted from the sub-semantic features to obtain the first keywords. The multi-head self-attention mechanism can capture the associated information of keywords in different contexts, enrich the semantic expression of keywords, and obtain more comprehensive and contextually relevant keyword semantic text. This enables a deeper and more accurate understanding of the user's style preference and provides strong support for determining the final style preference. For example, based on the sub-semantic features, the BERT language model extracts the first keywords such as "floral pattern" and "body-hugging". Through the multi-head self-attention mechanism, combined with the relevant context of the dress, keyword semantic text such as "vintage floral dress, body-hugging design highlights body curves" is obtained, which more comprehensively describes the dress style that the user is interested in. Finally, based on the preset text library, the style preference features of the keyword semantic text are extracted. The preset text library stores the definitions and descriptions of various styles. By matching with the keyword semantic text, the style preference features of the user can be quickly determined. This converts complex semantic information into specific style categories, enabling the e-commerce platform to provide more accurate search results and personalized recommendations for users according to these features.
[0024] In one embodiment, step S3 of performing semantic extraction on the search keyword to obtain an explicit semantic feature and obtaining an extended semantic feature according to the explicit semantic feature and the style preference feature includes: S301. Perform word segmentation on the search keyword to obtain multiple independent words of the search keyword, and convert the multiple independent words into word vector representations based on the Word2Vec model to obtain an independent word set vector, and define the semantic feature corresponding to the independent word set vector as the explicit semantic feature; S302. Map the style preference feature through a fully connected layer into a style semantic vector aligned with the dimension of the independent word set vector; S303. Compare and screen the independent word set vectors with the style semantic vectors based on the cosine similarity model to obtain independent word-style semantic subset vectors; S304. Vector splice the independent word-style semantic subset vectors and the style semantic vectors based on a preset time series to obtain a spliced vector; S305. Input the spliced vector into a preset bidirectional LSTM network for context semantic enhancement to obtain an enhanced spliced vector, and use the semantic features of the enhanced spliced vector as extended semantic features.
[0025] As described in the above steps S301 - S305, the present invention first performs word segmentation on the search keywords to obtain multiple independent words of the search keywords, and based on the Word2Vec model, converts the multiple independent words into word vector representations, obtaining an independent word set vector. The semantic features corresponding to the independent word set vector are defined as explicit semantic features. By word segmentation, complex search keywords are broken down into basic units, facilitating model understanding and processing. The Word2Vec model converts words into vectors, enabling the semantics of words to be presented in numerical form. Among them, Word2Vec is a word embedding technology based on neural networks, used to map words in natural language to low-dimensional dense vectors (usually 50 - 300 dimensions), such that words with similar semantics are close in the vector space. The independent word set vector integrates the semantic information of the keywords, and the explicit semantic features clarify the direct semantics of the keywords, providing a basis for subsequent semantic expansion. This helps the model accurately grasp the core content of the user's search intention and provides a basis for the preliminary screening of search results. Secondly, the style preference features are mapped through a fully connected layer into a style semantic vector aligned with the dimension of the independent word set vector. Style preference features usually have a specific representation form, and the fully connected layer can transform them to make their dimensions consistent with the independent word set vector. In this way, the style preference features can be compared and fused with the semantic vector of the keywords in the same space, facilitating the exploration of potential connections between the two, providing style-related information supplementation for expanding semantic features, and making the search results more in line with the user's personalized style requirements. For example, if through previous step analysis, the style preference feature of this user is "sports style with trendy elements", use the fully connected layer to map it into a style semantic vector with the same dimension as the above-mentioned independent word set vector. For example, encode the style preference of "sports style with trendy elements" into vector form to make its dimension consistent with the vectors corresponding to "fashionable" and "sports shoes". In subsequent steps, it will be better to combine the style preference with the semantics of the search keywords. By comparing and screening the independent word set vector and the style semantic vector based on the cosine similarity model, an independent word-style semantic subset vector is obtained. The cosine similarity model can measure the similarity degree between two vectors. Through comparison and screening, independent words with relatively high similarity to the style semantic vector and their related semantic parts can be found to form a subset vector. This step can screen out the part related to the user's style preference from the keyword semantics, highlight the semantic information matching the style, provide a more accurate direction for semantic expansion, reduce the interference of irrelevant information, and improve the matching degree between the search results and the user's style preference. For example, calculate the cosine similarity between the independent word set vector composed of "fashionable" and "sports shoes" and the style semantic vector of "sports style with trendy elements".It is found that the "trend" part in "fashion" and the "sports" attribute of "sports shoes" have a relatively high similarity with the style semantic vector, thus obtaining an independent word-style semantic subset vector, which contains the semantic vectors related to "fashion (trend part)" and "sports shoes (sports attribute)". During the search, the platform will pay more attention to the products with these semantic features, such as sports-style shoes with trendy design elements. Then, based on the preset time series, the independent word-style semantic subset vector and the style semantic vector are vector concatenated to obtain a concatenated vector. In this way, the preset time series provides an ordered way for vector concatenation, concatenating the subset vector and the style semantic vector in a specific order, which can integrate the semantic information of keywords and style preferences. The concatenated vector contains richer semantic content, covering both the part related to keywords and style and the overall style information, providing more comprehensive data for subsequent semantic enhancement, which helps to generate extended semantic features that better meet the user's needs. Finally, the concatenated vector is input into a preset bidirectional LSTM network for context semantic enhancement to obtain an enhanced concatenated vector, and the semantic features of the enhanced concatenated vector are used as the extended semantic features. In this way, the bidirectional LSTM network can learn the forward and backward information of the input sequence simultaneously and fully capture the context semantics. After processing the concatenated vector, the obtained enhanced concatenated vector has richer and more accurate semantics. The extended semantic features are generated based on this, which can better reflect the potential meaning of the user's search intention, enabling the search results to not only include directly relevant content but also be associated with context-related information, improving the comprehensiveness and accuracy of the search. For example: inputting the above concatenated vector into the bidirectional LSTM network, the network learns the connection between "fashionable sports shoes" and "sports style with trendy elements" in the context, such as the matching of trendy sports shoes in sports scenarios, and the styles of fashionable sports shoes suitable for different seasons. The obtained enhanced concatenated vector contains these rich context semantics, and its semantic features become the extended semantic features. During the search, based on these extended semantic features, the platform may recommend not only ordinary fashionable sports shoes but also the matching schemes of trendy sports shoes suitable for different sports scenarios and the styles of fashionable sports shoes popular in the current season.
[0026] In one embodiment, step S4 of obtaining the repurchase data information and the timeliness information according to the purchased product feature information and obtaining the purchase preference features according to the repurchase data information and the timeliness information includes: S401. Obtain the historical purchase timestamps and multiple product categories according to the purchased product feature information, and obtain the time intervals between two adjacent purchases under each same product category; S402. Mark the behaviors with multiple time intervals less than a preset threshold as repurchase events, and define the data information corresponding to the repurchase events as the repurchase data information; S403. Obtain a time sliding window according to the historical purchase timestamp, obtain a seasonal distribution interval according to the time sliding window, divide multiple said commodity categories into the seasonal distribution interval to obtain a season interval normalized commodity category, and use the time information of the season interval normalized commodity category in the seasonal distribution interval as timeliness information; S404. Perform timeliness screening on the repurchase data information according to the timeliness information to obtain screened repurchase data information, and perform encoding processing on the screened repurchase data information to obtain a screened repurchase data code; S405. Based on the Word2Vec model, convert the screened repurchase data code into a word vector to obtain a screened repurchase data code vector, and use the screened repurchase data code vector as a purchase preference feature.
[0027] As described in the above steps S401 - S405, the present invention first obtains the historical purchase timestamp and multiple product categories according to the purchase product feature information, and obtains the time intervals between two adjacent purchases under each same product category. In this way, by obtaining the historical purchase timestamp and product categories, the time context of the user's purchase behavior and the types of products involved can be clearly presented. Calculating the time intervals between two adjacent purchases under the same product category can reveal the differences in the purchase frequencies of different types of products by the user. This helps the e - commerce platform understand the user's consumption rhythm, which is crucial for judging the demand stability and repurchase potential of various products by the user. Then, the behaviors with multiple time intervals less than the preset threshold are marked as repurchase events, and the data information corresponding to the repurchase events is defined as repurchase data information. By setting the preset threshold to define repurchase events, the behaviors of users' repeated purchases can be accurately screened out. The repurchase data information centrally reflects the user's recognition and continuous demand for specific products, which is an important basis for judging the user's purchase preferences. The e - commerce platform can thereby understand which products are deeply loved by users, providing strong support for precise marketing and recommendations. Then, a time - sliding window is obtained according to the historical purchase timestamp, and a seasonal distribution interval is obtained according to the time - sliding window. Multiple product categories are divided into the seasonal distribution interval to obtain the season - interval - normalized product categories, and the time information of the season - interval - normalized product categories in the seasonal distribution interval is used as the timeliness information. By applying the time - sliding window and the seasonal distribution interval, the user's purchase behavior is combined with time - season factors. Among them, the time - sliding window is a technical means widely used in the processing of time - series data. By setting a window with a fixed length on the time series and allowing it to move step by step along the time axis to analyze the data. Its core principle lies in using the movement of the window to dynamically observe and analyze the local data characteristics in the time series. By dividing the product categories by season, the seasonal pattern of the user's purchase behavior can be found, that is, the demand differences for different product categories in different seasons. The acquisition of the timeliness information provides the e - commerce platform with more detailed user purchase time preferences, helping to recommend relevant products to users in the appropriate season and improving the timeliness and accuracy of recommendations. For example: Taking the purchase data of user A as an example, through the time - sliding window and season division, it is found that she has a higher purchase frequency of baby clothes in spring (March - May), and the product categories purchased mainly include thin coats, one - piece clothes, etc. These product categories are divided into the spring seasonal distribution interval and become the season - interval - normalized product categories. The time information of spring is the timeliness information related to these products.The platform can specifically recommend more clothing styles suitable for babies to wear in spring to user A during spring. Then, according to the timeliness information, the repurchase data information is screened for timeliness to obtain the screened repurchase data information. The screened repurchase data information is encoded to obtain the screened repurchase data code. In this way, screening the repurchase data information using timeliness information can further focus on the commodity data with repurchase behavior within a specific season and exclude the interfering data that does not meet the seasonal requirements. Encoding processing converts the screened repurchase data into a format suitable for model processing, facilitating subsequent in-depth analysis and feature extraction, enabling the obtained purchase preference features to more accurately reflect the actual purchase needs of users in different seasons. Secondly, based on the Word2Vec model, the screened repurchase data code is converted into a word vector to obtain the screened repurchase data code vector. The screened repurchase data code vector is used as the purchase preference feature. By converting the screened repurchase data code into a word vector through the Word2Vec model, the semantic information of the data can be represented in vector form, facilitating the model's calculation and analysis. The screened repurchase data code vector integrates the repurchase commodity information of users within a specific season and comprehensively reflects the purchase preference features of users. The e-commerce platform can, based on these vector features, more accurately provide commodity recommendations that match the users' purchase preferences and optimize search results, improving users' shopping satisfaction and the platform's operation efficiency. For example, the screened repurchase data code of user A in spring is converted into a vector using the Word2Vec model. Suppose the code vector of milk powder is represented as [0.1, 0.3, 0.5,...], the code vector of diapers is [0.2, 0.4, 0.6,...], and the code vector of baby thin coats is [0.7, 0.1, 0.8,...]. These vectors are combined as the purchase preference features of user A in spring. When the platform searches for relevant commodities for the user, based on these vector features, it preferentially displays commodities such as milk powder brands, diaper styles, and baby thin coat styles that match user A's purchase preferences.
[0028] In one embodiment, step S5 of multiplying and quantizing the extended semantic feature and the purchase preference feature in segments to obtain the compressed semantic association feature includes: S501. Mapping the extended semantic feature to an extended semantic feature vector; S502. Mapping the purchase preference feature to a purchase preference feature vector, and splicing the extended semantic feature vector and the purchase preference feature vector to obtain a joint vector; S503. Dividing the joint vector into multiple sub-vector segments according to a preset length, and clustering each sub-vector segment based on K-means clustering to obtain multiple clustering centroids, where each clustering centroid represents an index; S504. Map multiple of the cluster centroids to multiple cluster centroid vectors, perform product quantization on the multiple cluster centroid vectors to obtain a mass point vector value, and use the semantics corresponding to the mass point vector value as the compressed semantic association feature.
[0029] As described in the above steps S501 - S504, in the fields of deep learning and data processing, vectors are a commonly used data representation form. Converting semantic features into vectors conforms to the way computers process data, enabling complex semantic information to be transmitted and processed more efficiently in the model. In the present invention, the extended semantic features are first mapped to extended semantic feature vectors, converting the extended semantic features into vector form, which allows the computer to understand and process this semantic information in a numerical way. The vector form facilitates various mathematical operations and analyses. After quantifying the semantic features, comparison, fusion, and other operations can be performed with other vectors in subsequent steps, thereby mining deeper semantic associations and improving the accuracy of understanding the user's search intent. Secondly, the purchase preference features are mapped to purchase preference feature vectors, and the extended semantic feature vectors and the purchase preference feature vectors are concatenated to obtain a joint vector. Converting the purchase preference features into vectors and concatenating them with the extended semantic feature vectors realizes the fusion of information from two aspects: the user's search intent and purchase habits. The joint vector contains more comprehensive user demand information, including not only the semantic extensions related to the current search keywords but also the preferences reflected by the user's past purchase behaviors, providing a rich data basis for more accurately screening and matching search results in the future. Then, the joint vector is divided into multiple sub - vector segments according to a preset length, and each sub - vector segment is clustered based on K - means clustering to obtain multiple cluster centroids. Among them, each cluster centroid represents an index. Dividing the joint vector into sub - vector segments in this way can reduce the data dimension and computational complexity. K - means clustering can group similar sub - vector segments together and find their cluster centroids. Each cluster centroid, as an index, can represent a class of vectors with similar characteristics, which helps classify and quickly retrieve a large amount of data. In this way, data with similar user demands can be grouped into one category, improving the search efficiency and accuracy. For example: if the length of the joint vector is 10 and it is preset to be divided into sub - vector segments with a length of 3, 3 sub - vector segments are obtained (the last segment with less than 3 elements can be filled or specially processed). Using the K - means clustering algorithm for these sub - vector segments, assuming K is 3 (i.e., divided into 3 categories), after clustering, 3 cluster centroids are obtained, which are [0.7, 0.6, 0.8], [0.5, 0.4, 0.3], and [0.9, 0.8, 0.7] respectively. These 3 cluster centroids represent different categories of user demand characteristics. For example, the first cluster centroid represents a category of demands with higher requirements for the functions and performance of sports watches. Each cluster centroid is stored in the database as an index.When there is a new search request, the commodity data that matches it can be quickly found through these indexes. Finally, map multiple said clustering centroids into multiple said clustering centroid vectors, and perform product quantization on the multiple said clustering centroid vectors to obtain a particle vector value. Use the semantics corresponding to the particle vector value as the compressed semantic association feature. In this way, mapping the clustering centroids into vectors again and performing product quantization further compresses the data dimension, reduces the storage space and the amount of calculation. The compressed semantic association feature retains the core semantic association information of the data, improves the running efficiency of the system while ensuring the search accuracy. Such compressed features can be used more efficiently for subsequent index construction and search matching, making the search process faster and more accurate.
[0030] In one embodiment, the step S5 of simultaneously performing double-layer compact indexing on the explicit semantic feature and the compressed semantic association feature to obtain an explicit result and an association result includes: S505. Construct a first-layer index based on the explicit feature inverted index for the explicit semantic feature, and map the first-layer index to a hash bucket based on Locality-Sensitive Hashing (LSH) to obtain a first-layer hash index. Use the result corresponding to the first-layer hash index as the explicit result; S506. Construct a second-layer index for the compressed semantic association feature based on the compressed feature map index, and construct the second-layer index into a multi-layer graph index based on the Hierarchical Navigable Small World (HNSW) graph algorithm, where each layer of the graph index corresponds to a commodity identifier; S507. Use the result corresponding to the multi-layer graph index as the association result.
[0031] As described in the above steps S505 - S507, the present invention first constructs the first - layer index based on the explicit semantic features using an explicit - feature inverted index, and maps the first - layer index to hash buckets based on Locality - Sensitive Hashing (LSH) to obtain the first - layer hash index. The result corresponding to the first - layer hash index is used as the explicit result. In this way, the explicit - feature inverted index can quickly locate the documents (product data) containing specific explicit semantic features, improving the retrieval efficiency. Locality - Sensitive Hashing (LSH) can map similar vectors to the same hash bucket, reducing the comparison range and further accelerating the retrieval speed. The explicit result provides users with basic search results directly related to the search keywords, meeting the basic needs of users and enabling them to quickly obtain the most direct and relevant product information. Then, based on the compressed - feature map index, the second - layer index is constructed for the compressed semantic - association features, and a multi - layer graph index is constructed for the second - layer index based on the Hierarchical Navigable Small World (HNSW) graph algorithm. Each layer of the graph index corresponds to a product identifier. By constructing the compressed - feature map index for the compressed semantic - association features, these complex semantic information can be effectively organized and stored. The multi - layer graph index constructed by the HNSW graph algorithm can achieve efficient approximate nearest - neighbor search. In this way, the system can quickly locate relevant product identifiers based on the potential needs of users (compressed semantic - association features), discover products potentially associated with the user's search intent, enrich the search results, and improve the accuracy and comprehensiveness of the search. Finally, the result corresponding to the multi - layer graph index is used as the associated result. By presenting the products corresponding to the multi - layer graph index as the associated results to users, the search results can be supplemented and expanded. These associated results are based on the potential needs and behavior patterns of users, providing users with more products that they may be interested in but have not explicitly searched for, enhancing the user shopping experience, increasing the exposure opportunities and sales possibilities of products.
[0032] This application also provides a private - domain e - commerce data search system based on big data, including: A first acquisition module, configured to acquire user search feature information of a private - domain e - commerce platform, where the user search feature information includes behavior - feature information and purchased - product feature information; A second acquisition module, configured to acquire browsing - feature information and search keywords according to the behavior - feature information, acquire the click - through times of multiple browsed products and the click - through stay times of multiple browsed products according to the browsing - feature information, and acquire style - preference features according to the click - through times of multiple browsed products and the click - through stay times of multiple browsed products; An extraction module, configured to perform semantic extraction on the search keywords to obtain explicit semantic features, and obtain extended semantic features according to the explicit semantic features and style - preference features; A third acquisition module, configured to acquire repurchase - data information and timeliness information according to the purchased - product feature information, and acquire purchase - preference features according to the repurchase - data information and timeliness information; A fourth acquisition module, configured to perform segmented product quantization compression on the extended semantic features and the purchase preference features to obtain compressed semantic association features, and perform simultaneous double-layer compact indexing on the explicit semantic features and the compressed semantic association features to obtain an explicit result and an association result; A presentation module, configured to visually present the private domain e-commerce search results according to the explicit result and the association result.
[0033] In one embodiment, the second acquisition module includes: A first sorting unit, configured to perform a descending sort on the click times of multiple browsing products to obtain a browsing product click time sorting table; A second sorting unit, configured to perform a descending sort on the browsing click stay times of multiple products to obtain a browsing click stay time sorting table; A first acquisition unit, configured to acquire a first product corresponding to the first browsing product click time according to the browsing product click time sorting table; A second acquisition unit, configured to acquire a second product corresponding to the first browsing click stay time sorting table according to the browsing click stay time sorting table, and perform dynamic load coefficient constraint on the second product and the first product to obtain an equilibrium clustering cluster set, and perform centroid screening on the equilibrium clustering cluster set based on the K-means clustering algorithm to obtain a product centroid; A screening unit, configured to perform fine-grained semantic feature screening on the product centroid based on an attention mechanism to obtain multiple fine-grained semantic features; A first extraction unit, configured to perform keyword extraction on multiple fine-grained semantic features based on a BERT language model to obtain first keywords, and perform context association capture on the first keywords based on a multi-head self-attention mechanism to obtain keyword semantic texts; A second extraction unit, configured to extract style preference features of the keyword semantic texts based on a preset text library.
[0034] This application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0035] This application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0036] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be obtained in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0037] It should be noted that in this text, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that includes a series of elements includes not only those elements but also other elements not explicitly listed, or further includes elements inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method that includes such element.
[0038] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made using the specifications and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included within the patent protection scope of the present invention.
Claims
1. A private domain e-commerce data search method based on big data, characterized in that: include: Obtain user search feature information of the private domain e-commerce platform, wherein the user search feature information includes behavior feature information and purchase product feature information; Obtain browsing characteristic information and search keywords according to the behavior characteristic information, obtain multiple browsed product click times and multiple browse click dwell times according to the browsing characteristic information, and obtain style preference characteristics according to the multiple browsed product click times and multiple browse click dwell times; Performing semantic extraction on the search keyword to obtain explicit semantic features, and obtaining extended semantic features based on the explicit semantic features and style preference features; Acquire repurchase data information and timeliness information according to the purchased commodity characteristic information, and acquire purchase preference characteristics according to the repurchase data information and timeliness information; The extended semantic features and the purchase preference features are subjected to piecewise product quantization compression to obtain compressed semantic association features, and the explicit semantic features and the compressed semantic association features are subjected to dual-layer compact indexing to obtain explicit results and association results; The private domain e-commerce search results are visualized according to the explicit results and the associated results.
2. The method for searching private domain e-commerce data based on big data according to claim 1 is characterized in that: The step of acquiring style preference features according to the plurality of browsed product click times and the plurality of browse click dwell times comprises: Sorting the number of clicks on the browsed products in descending order to obtain a ranking table of the number of clicks on the browsed products; Sort the plurality of browsing click dwelling times in descending order to obtain a browsing click dwelling time sorting table; Obtaining the first product corresponding to the number of clicks of the first browsed product according to the browsed product click count ranking table; According to the browsing click and stay time ranking table, the second product corresponding to the first browsing click and stay time ranking table is obtained, and the second product and the first product are subjected to dynamic load coefficient constraints to obtain a balanced clustering cluster set, and the balanced clustering cluster set is subjected to centroid screening based on the K-means clustering algorithm to obtain the product centroid; Based on the attention mechanism, segmentation semantic features are screened for the centroid of the product to obtain multiple segmentation semantic features; Perform keyword extraction on the plurality of segmented semantic features based on the BERT language model to obtain a first keyword, and perform context association capture on the first keyword based on a multi-head self-attention mechanism to obtain a keyword semantic text; The style preference features of the keyword semantic text are extracted based on a preset text library.
3. The method for searching private domain e-commerce data based on big data according to claim 1 is characterized in that: The step of performing semantic extraction on the search keyword to obtain explicit semantic features, and obtaining extended semantic features according to the explicit semantic features and style preference features includes: Performing word segmentation processing on the search keyword to obtain multiple independent words of the search keyword, and converting the multiple independent words into word vectors based on the Word2Vec model to obtain independent word set vectors, and defining the semantic features corresponding to the independent word set vectors as explicit semantic features; Mapping the style preference feature into a style semantic vector aligned with the dimension of the independent word set vector through a fully connected layer; Based on the cosine similarity model, the independent word set vector is compared and screened with the style semantic vector to obtain an independent word-style semantic subset vector; Based on a preset time sequence, the independent word-style semantic subset vector and the style semantic vector are concatenated to obtain a concatenated vector; The splicing vector is input into a preset bidirectional LSTM network for contextual semantic enhancement to obtain an enhanced splicing vector, and the semantic features of the enhanced splicing vector are used as extended semantic features.
4. The method for searching private domain e-commerce data based on big data according to claim 1, characterized in that: The step of acquiring repurchase data information and timeliness information according to the purchased commodity characteristic information, and acquiring purchase preference characteristics according to the repurchase data information and timeliness information includes: Obtaining historical purchase timestamps and multiple commodity categories according to the purchased commodity feature information, and obtaining the time interval between two adjacent purchases under each of the same commodity categories; Marking a plurality of behaviors whose time intervals are less than a preset threshold as repurchase events, and defining the data information corresponding to the repurchase events as repurchase data information; Obtaining a time sliding window according to the historical purchase timestamp, and obtaining a seasonal distribution interval according to the time sliding window, dividing the plurality of commodity categories into seasonal distribution intervals, obtaining seasonal interval naturalized commodity categories, and using time information of the seasonal interval naturalized commodity categories in the seasonal distribution interval as timeliness information; Performing time screening on the repurchase data information according to the time information to obtain screened repurchase data information, and performing encoding processing on the screened repurchase data information to obtain a screened repurchase data code; Based on the Word2Vec model, the filtered repurchase data encoding is converted into a word vector to obtain the filtered repurchase data encoding vector, and the filtered repurchase data encoding vector is used as a purchase preference feature.
5. The method for searching private domain e-commerce data based on big data according to claim 1 is characterized in that: The step of performing piecewise product quantization compression on the extended semantic features and the purchase preference features to obtain compressed semantic association features comprises: Mapping the extended semantic feature into an extended semantic feature vector; Mapping the purchase preference feature into a purchase preference feature vector, and concatenating the extended semantic feature vector and the purchase preference feature vector to obtain a joint vector; Divide the joint vector into a plurality of sub-vector segments according to a preset length, and cluster each sub-vector segment based on K-means clustering to obtain a plurality of cluster centroids, wherein each cluster centroid represents an index; The plurality of cluster centroids are mapped into the plurality of cluster centroid vectors, and the plurality of cluster centroid vectors are multiplied and quantized to obtain particle vector values, and the semantics corresponding to the particle vector values are used as compressed semantic association features.
6. The method for searching private domain e-commerce data based on big data according to claim 1 is characterized in that: The step of simultaneously performing double-layer compact indexing on the explicit semantic features and the compressed semantic association features to obtain explicit results and association results includes: The explicit semantic feature is constructed based on the explicit feature inverted index to obtain a first-level index, and the first-level index is mapped to a hash bucket based on local sensitive hashing (LSH) to obtain a hash first-level index, and a result corresponding to the hash first-level index is used as an explicit result; Based on the compressed feature graph index, the compressed semantic association features are used to construct a second-layer index, and based on a hierarchical navigable small world (HNSW) graph algorithm, the second-layer index is constructed into a multi-layer graph index, wherein each layer of the graph index corresponds to a product identifier; The result corresponding to the multi-layer graph index is used as the association result.
7. A private domain e-commerce data search system based on big data, characterized in that: include: The first acquisition module is used to obtain user search feature information of the private domain e-commerce platform, wherein the user search feature information includes behavior feature information and purchase commodity feature information; A second acquisition module is used to acquire browsing characteristic information and search keywords according to the behavior characteristic information, and acquire multiple browsed product click times and multiple browse click dwell times according to the browsing characteristic information, and acquire style preference characteristics according to the multiple browsed product click times and multiple browse click dwell times; An extraction module, used for performing semantic extraction on the search keyword to obtain explicit semantic features, and obtaining extended semantic features based on the explicit semantic features and style preference features; A third acquisition module is used to acquire repurchase data information and timeliness information according to the purchased commodity characteristic information, and acquire purchase preference characteristics according to the repurchase data information and timeliness information; The fourth acquisition module is used to perform piecewise product quantization compression on the extended semantic features and the purchase preference features to obtain compressed semantic association features, and to perform double-layer compact indexing on the explicit semantic features and the compressed semantic association features to obtain explicit results and association results; A presentation module is used to visually present the private domain e-commerce search results based on the explicit results and the associated results.
8. A private domain e-commerce data search system based on big data according to claim 7, characterized in that: The second acquisition module includes: A first sorting unit is used to sort the click counts of the browsed products in descending order to obtain a sorting table of click counts of browsed products; A second sorting unit is used to sort the plurality of browsing click dwelling times in descending order to obtain a browsing click dwelling time sorting table; A first acquisition unit, configured to acquire a first product corresponding to a first number of clicks on a first browsed product according to the browsed product click count ranking table; A second acquisition unit is used to acquire a second commodity corresponding to the first-ranked browsing click and stay time ranking table according to the browsing click and stay time ranking table, and subject the second commodity and the first commodity to dynamic load coefficient constraints to obtain a balanced clustering cluster set, and perform centroid screening on the balanced clustering cluster set based on a K-means clustering algorithm to obtain a commodity centroid; A screening unit, configured to screen the centroid of the product for segmented semantic features based on an attention mechanism to obtain a plurality of segmented semantic features; A first extraction unit is used to perform keyword extraction on the plurality of segmented semantic features based on a BERT language model to obtain a first keyword, and to capture context association of the first keyword based on a multi-head self-attention mechanism to obtain a keyword semantic text; The second extraction unit is used to extract the style preference features of the keyword semantic text based on a preset text library.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Personalized image browsing and recommending method based on labelling semantics and system thereof
CN102663010A
Searching shopping method based on pictures
CN104765891A
Commodity recommendation method and system based on combination of twin Transform model and knowledge graph
CN115934967A
Commodity intelligent recommendation system and method based on data mining technology
CN119006119A
Intelligent query semantic understanding method based on information geometry and Riemannian manifold
CN119829740A
Cited By
Electronic commerce data processing method and system
CN120471650A
Feed raw material traceable management system
CN121146798A