Commodity recommendation method based on big data

By combining attribute similarity and semantic similarity calculations and dynamically adjusting multi-layered interest boundaries, the problem of insufficient utilization of product data diversity and unstructured data in traditional recommendation systems is solved, enabling more accurate and diversified product recommendations, improving user experience and platform development.

CN121961697APending Publication Date: 2026-05-01BEIJING THE GREAT WALL AGEL ECOMMERCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING THE GREAT WALL AGEL ECOMMERCE CO LTD
Filing Date
2026-01-22
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional recommendation systems neglect the diversity and complexity of product data, especially unstructured data, which is difficult to utilize effectively. This leads to insufficient recommendation accuracy and can narrow users' horizons, inhibiting their ability to explore new interests.

Method used

By employing a hybrid calculation of attribute similarity and semantic similarity, and combining information credibility score as a weighting factor, the recommendation strategy is dynamically adjusted through the construction of a multi-layer interest boundary and negative feedback boundary correction mechanism. The collaborative control of exploration gain coefficient and boundary chain is introduced to achieve accurate, reliable and adaptive product recommendation.

Benefits of technology

It improves the accuracy and diversity of recommendations, dynamically adjusts recommendation strategies to meet users' current interests, guides users to discover new interests, enhances user experience and loyalty, reduces interference from low-credibility products, and minimizes recommendation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961697A_ABST
    Figure CN121961697A_ABST
Patent Text Reader

Abstract

The invention discloses a commodity recommendation method based on big data, and the method comprises the steps: collecting data, recognizing the credibility of the data, carrying out the statement analysis of the data based on a big model, and generating an information credibility score; calculating attribute similarity and semantic similarity, calculating basic similarity, and correcting the basic similarity to obtain final similarity; a weight mapping table is set, the system performs automatic query according to commodity categories, obtains an optimal weight combination and performs descending sort according to final similarity to obtain a similar commodity recommendation table, and a large model is used for generating an explanatory reason for a recommendation result in the similar commodity recommendation table; when similar commodities are insufficient or are new commodities, a popularity substitution strategy is used, mixed calculation of attribute similarity and semantic similarity is adopted, recommendation accuracy is improved, information credibility scores are introduced to serve as weight factors, similarity calculation is more accurate, interference of low-credibility commodities on recommendation results is reduced, and recommendation quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic product recommendation technology, and in particular to a product recommendation method based on big data. Background Technology

[0002] With the booming development of the e-commerce industry, users spend a considerable amount of time on product detail pages and have high purchase conversion potential. Therefore, product detail pages have become an important platform for recommending similar products, aiming to meet users' comparison needs while browsing products, helping them discover more potentially interesting items, thereby improving the user shopping experience and the platform's sales performance.

[0003] Traditional recommendation systems often overlook the diversity and complexity of product data. Product information on e-commerce platforms comes from a wide range of sources, including but not limited to product descriptions, user reviews, images, and videos. These data vary significantly in format, quality, and reliability. Effectively integrating and utilizing this multi-source, heterogeneous data has become a major challenge in improving recommendation accuracy. In particular, unstructured data such as images and videos contain rich product feature information, but traditional methods often struggle to directly utilize this information for recommendations. Summary of the Invention

[0004] This application provides a product recommendation method based on big data, which solves the problem that traditional product recommendation ignores the diversity and complexity of product data. By adopting a hybrid calculation of attribute similarity and semantic similarity, the recommendation accuracy is improved compared with a single algorithm. By introducing information credibility score as a weighting factor, the similarity calculation is more accurate, reducing the interference of low credibility products on the recommendation results and improving the recommendation quality.

[0005] This application provides a product recommendation method based on big data, including: S101 collects users' historical interaction behavior data, identifies the core area based on the historical interaction behavior data, delineates the exploration area based on the core area, identifies the peripheral area based on the exploration area, and constructs a multi-layered interest boundary based on the core area, exploration area, and peripheral area; S102, collect product data, calculate attribute similarity and semantic similarity based on product data, calculate basic similarity through attribute similarity and semantic similarity, correct the basic similarity to obtain the final similarity, obtain the basic fit layer score based on the final similarity, and calculate the exploration gain coefficient layer by layer based on the multi-layer interest boundary. S103 generates an interest recommendation list based on the exploration gain coefficient, tracks user behavior data in real time, and updates the multi-layer interest boundaries.

[0006] Preferably, the steps for calculating the final similarity are as follows: converting the standardized attribute information into attribute vectors and calculating attribute similarity; generating text vectors from the standardized description text using an algorithm and calculating semantic similarity; calculating basic similarity based on the weighted combination of attribute similarity and semantic similarity, combined with the information credibility score; and correcting the basic similarity to obtain the final similarity.

[0007] Preferably, the steps in S101 of identifying the core area based on historical interaction behavior data, delineating the exploration area based on the core area, and identifying the outer area based on the exploration area are as follows: identifying the core area, exploration area, and outer area based on the calculated Euclidean distance, if d≤ d represents the Euclidean distance, and the core area is centered on the user's core interest point. A spherical region with radius ; if d≤ If the item falls into the exploration zone, the exploration zone is defined by its inner diameter being the radius of the core area. The outer diameter is A ring-shaped region; if d If the item falls into the outer zone, the outer zone is the area outside the exploration zone. All vector spaces other than that.

[0008] Preferably, the method for constructing multi-layer interest boundaries in S101 is as follows: by identifying high-frequency user interaction behaviors, core interest points are calculated, and a core radius is defined by the core interest points; the core radius is used as the inner diameter of the exploration area, and the outer diameter of the exploration area is obtained through the core interest points, thereby defining the exploration area; all vector spaces outside the outer diameter of the exploration area are defined as the outer region.

[0009] Preferably, generating the recommendation list based on the exploration gain coefficient further includes: S301, collect negative feedback signals, construct a set of negative feedback commodity vectors based on the negative feedback signals, calculate negative feedback cluster centers based on the set of negative feedback commodity vectors, calculate the distance from candidate commodity vectors to negative feedback cluster centers, and combine the calculated distances with a function to construct a repulsion vector field; S302, calculate the minimum distance based on the center of the aversion interest cluster, set a safe distance threshold based on the minimum distance, delineate the interest safety zone in reverse by the center of the aversion interest cluster and the safe distance threshold, and correct the multi-layer interest boundary based on the interest safety zone; S303, based on the repulsion vector field, corrects the exploration gain coefficient and uses the corrected exploration gain coefficient to generate a safe recommendation list.

[0010] Preferably, the step of collecting negative feedback signals involves collecting explicit feedback and implicit feedback, with the explicit feedback having a higher intensity than the implicit feedback.

[0011] Preferably, the method further includes: S401 collects external event data, monitors the intensity of external events, calculates the interest drift index based on historical interaction data, generates boundary chain nodes based on the intensity of external events and the interest drift index, and constructs the boundary chain; S402, use a prediction model to generate interest nodes, and build a recommendation chain based on the path vectors of the boundary chain nodes and the interest nodes; S403, the boundary chain and the recommendation chain are controlled collaboratively, a recommendation score is calculated, an intelligent recommendation list is generated based on the recommendation score, and the boundary chain and the recommendation chain are cross-adjusted. The cross-adjustment is that the boundary chain adjusts the recommendation chain, and the path strength of the recommendation chain adjusts the core area radius of the boundary chain.

[0012] One or more technical solutions provided in this application have at least the following technical effects or advantages: the hybrid calculation of attribute similarity and semantic similarity improves the recommendation accuracy compared to a single algorithm; by introducing information credibility score as a weight factor, the similarity calculation is more accurate, reducing the interference of low-credibility products on the recommendation results and improving the recommendation quality; the dynamic weight allocation mechanism adjusts the weight according to the characteristics of the product category, making the recommendation strategy more accurate and flexible; and a recommendation system that integrates product information credibility assessment, dynamic weight allocation and negative feedback loop is constructed to achieve accurate, reliable and adaptive recommendation of candidate products, thereby improving user trust and purchasing efficiency. By introducing an exploration gain coefficient and constructing a dynamic multi-layered interest boundary for users, the system increases the diversity and exploratory nature of recommendations while ensuring their relevance. It appropriately reduces the weight of core area products through a decay factor to allow room for exploration; it enhances the ranking of exploration area products through boundary expansion layer scoring and innovation premium, proactively guiding users to expand their interest boundaries; and it strongly suppresses the recommendation of peripheral area products through a minimal constant to avoid harming the user experience. This ensures that the recommendation results not only satisfy users' current interests but also guide them to discover new points of interest, enhancing the diversity and long-term value of the user experience. This helps increase user satisfaction and loyalty to the recommendation system, promoting the platform's long-term development. While ensuring the basic relevance of the recommendation results, it proactively and reasonably recommends products that can expand users' interest boundaries, enhancing the diversity and long-term value of the user experience. By using a negative feedback boundary correction mechanism, we can proactively avoid product features that users clearly dislike, reduce recommendation errors, and use dislike to define the boundaries of like. This provides a more refined and comprehensive understanding of user interest models, transforming negative feedback data from a passive blacklist tool into a precise probe that actively defines the user's preference space, thus enabling more accurate recommendations for products of interest. In emergency scenarios, dual-chain collaboration makes recommendation scores more accurate. Through boundary chains and collaborative controllers, interest boundaries and recommendation paths can be adjusted in real time, providing a more humanized, coherent, and accurate e-commerce recommendation solution, moving from static division to dynamic guidance. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating a product recommendation method based on big data according to the present invention. Figure 2 A schematic diagram illustrating the process of generating a product recommendation list for this invention; Figure 3 A schematic diagram illustrating the process of constructing the repulsion vector field for this invention; Figure 4 This is a schematic diagram illustrating the process of constructing the boundary chain and recommendation chain in this invention. Detailed Implementation

[0014] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.

[0015] It should be noted that the terms "vertical," "horizontal," "up," "down," "left," "right," and similar expressions used in this article are for illustrative purposes only and do not represent the only possible implementation.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0017] Example 1: Traditional recommendation methods are prone to over-optimizing short-term click-through rates, leading to information cocoons. The system will continuously recommend similar products that users already know and prefer. Although this is safe, it narrows the user's horizons and inhibits their ability to explore new interests, which will lead to user fatigue in the long run. Figure 1 This is a flowchart illustrating a product recommendation method based on big data according to an embodiment of the present invention, including: S101 collects users' historical interaction behavior data, identifies the core area based on historical interaction behavior, delineates the exploration area based on the core area, identifies the peripheral area based on the exploration area, and constructs a multi-layered interest boundary based on the core area, exploration area, and peripheral area. Specifically, all user interaction records from the past 90 days are collected. These records include clicks, browsing timeouts, adding to cart, and purchases. Clicks indicate user interest in a product; browsing timeouts suggest prolonged viewing of product content, potentially indicating strong interest; adding to cart shows an intention to purchase the product; and purchases directly reflect high user satisfaction. A large model (ChatGLM) is used, with each product interacted with as input. This model maps each product to a unified vector space, resulting in a set of historical product interaction vectors. A browsing time threshold is set to filter the historical product interaction vector set. Interactions exceeding the threshold or reaching a certain page depth are identified as high-frequency interactions. These high-frequency interactions constitute the core product set. For all vectors in the core product set, a weighted average vector is calculated to obtain the core interest points. ,in, The core interest vector is a weighted average of all product vectors in the core product set, representing the position of the user's core interest in the vector space. m is the total number of product vectors in the core product set, and j is the index of the product vector in the core product set. Let be the vector of the j-th product in the core product set. It is a vector representing the characteristics of that product in a unified vector space. For the j-th item vector The corresponding interaction behavior weights are determined based on the user's interaction behavior type. The distances from all vectors in the core product set to the core interest point are calculated, and the average of these distances is taken. 1.5 times the average is used as the radius of the core area, which is a circular region centered on the core interest point with a radius equal to 1.5 times the average. The exploration area is an annular region, including an inner and outer diameter. The inner diameter of the exploration area is directly taken as the outer edge of the core area, i.e., the value of the core area radius. For the outer diameter of the exploration area, the maximum distance from all historical interactive product vectors to the core interest point is calculated, and 1.2 times the maximum distance is used as the outer diameter of the exploration area. All vector spaces outside the outer diameter of the exploration area are defined as the outer perimeter. The outer perimeter contains product vector information that is relatively peripheral and weakly associated with the core interest in the user's historical interactions. Based on the core area, exploration area, and outer perimeter, a three-layer interest boundary is constructed.

[0018] S102, the basic fit layer score is obtained based on the final similarity, and the exploration gain coefficient is calculated layer by layer on the basis of the multi-layer interest boundary; Furthermore, based on the calculated final similarity, the final similarity is normalized using max-min normalization: ,in, The basic fit layer score is used, where x represents the final similarity score. These are the minimum and maximum similarity values ​​among all candidate products, respectively. The normalized result serves as the base fit layer score for this product. Candidate products refer to all products to be recommended. Features of candidate products are extracted, quantized, and then combined in order to form a candidate product vector. The Euclidean distance from the candidate product vector to the user's core interest point is calculated using the following formula: ,in, Let n be the Euclidean distance, and n be the dimension of the vector. Let i be the value of the i-th dimension of the candidate product vector. Given the value of the i-th dimension of the user's core interest, determine which region the product falls into based on the calculated Euclidean distance. The specific determination method is as follows: If d≤ If the product falls into the core area, the core area is centered on the user's core interests. A spherical region with radius r represents a set of products highly relevant to the user's core interests; if d≤ If the item falls into the exploration zone, the exploration zone is a zone with an inner diameter of [missing information]. The outer diameter is In the annular region, calculate the distance from the product to the inner boundary of the core area: =d- The products in the exploration area are related to users' core interests, but also bring new information; if d If the item falls into the outer zone, the outer zone is the area outside the exploration zone. All vector spaces outside the core area represent the set of products with weaker correlation to user interests; the exploration gain coefficient is calculated based on the different regions where the product falls. When a product falls into the core area, the formula for calculating the exploration gain coefficient is: ,in, For the exploration gain coefficient of the core area, Based on the basic fit layer score, The core area reward factor is used; when a product falls into the exploration area, a boundary expansion layer score is calculated. This boundary expansion layer assesses the product's potential to explore new user interests by quantifying its distance from the core area boundary. Products farther from the core area boundary have higher exploration value represented by their boundary expansion layer. The formula for calculating the boundary expansion layer score is: ,in, Scoring of the boundary extension layer The exploration gain coefficient of the exploration zone is calculated based on the boundary expansion layer score and the basic fit layer score, given the distance from the goods falling in the exploration zone to the boundary boundary of the core zone. The formula is as follows: ,in, This represents the exploration gain coefficient for the exploration zone. Based on the basic fit layer score, Scoring of the boundary extension layer To explore the incentive coefficient, To explore the attenuation factor, this embodiment... =1.2, a coefficient greater than 1, is used to further adjust the recommendation priority of goods in the exploration zone. Its function is to strengthen the system's encouragement of exploration behavior and highlight the recommendation value of goods in the exploration zone. Meanwhile, Adjustments are made based on actual recommendation performance and business needs to balance the diversity and relevance of recommended products; when a product falls into the outer zone, the formula for calculating the exploration gain coefficient is: ,in, For the exploration gain coefficient of the outer region, Based on the basic fit layer score, ϵ=0.001 is a very small constant. The products in the outer area are completely unrelated to the user's interests or have a very weak correlation. In order to avoid these products appearing in the recommendations and damaging the user experience, their recommendations are strongly suppressed by multiplying them by a very small coefficient ϵ=0.001, so that the priority of these products in the recommendations is almost zero.

[0019] S103 generates an interest recommendation list based on the exploration gain coefficient, tracks user behavior data in real time, and updates the multi-layer interest boundaries; Specifically, all products in the candidate product list are reordered from high to low according to their exploration gain coefficients. The top few products after reordering are used to form an interest recommendation list. For each product in the interest list, especially those exploration zone products whose ranking has improved due to high exploration gain coefficients, a recommendation reason can be generated using a large model (ChatGLM). The user's subsequent behavior towards the recommended products is tracked in real time. If the user frequently clicks on exploration zone products, the exploration encouragement coefficient can be appropriately increased; if the user has a lot of negative feedback on exploration zone products, the exploration encouragement coefficient is temporarily decreased. The user's new interaction behavior will be integrated into step S101 in real time to update their core interest points and the boundaries of each region, realizing the dynamic evolution of interest boundaries.

[0020] The technical solutions described in the embodiments of this application above have at least the following technical effects or advantages: By introducing an exploration gain coefficient and constructing a dynamic multi-layered interest boundary for users, the diversity and exploratory nature of recommendations are increased while ensuring recommendation relevance. The core area products are appropriately weighted by a decay factor, leaving room for exploration. The ranking of products in the exploration area is improved through boundary expansion layer scoring and innovation premium, actively guiding users to expand their interest boundaries. The recommendation of peripheral area products is strongly suppressed by a minimal constant, avoiding damage to the user experience. This ensures that the recommendation results not only satisfy the user's current interests but also guide the user to discover new points of interest, enhancing the diversity and long-term value of the user experience. This helps increase user satisfaction and loyalty to the recommendation system, promoting the long-term development of the platform. While ensuring the basic relevance of the recommendation results, it proactively and reasonably recommends products that can expand the user's interest boundaries, enhancing the diversity and long-term value of the user experience. Example 2: As... Figure 2 As shown, the product recommendation method further includes: S201, Collect data and identify the credibility of the data, and preprocess the collected data; Specifically, raw data is collected from multiple sources, including product information tables, product description tables, and product attribute tables. Basic information such as product ID, name, and category is collected from the product information table; text descriptions, selling points, and images are collected from the product description table; and specifications, parameters, and materials are collected from the product attribute table. Text data is standardized by removing garbled characters and special symbols, and a unified encoding format is applied. For image data, image URLs or storage paths are recorded, and image accessibility is checked. Information scattered across different tables is linked using the product ID to form a complete product record. The source type of the collected images is defined, including official rendered images, merchant-taken images, and images from unknown sources. The source is initially determined by the image filename, URL path, or watermark information. For example, an image filename containing "official_render" indicates an official rendered image, and an image URL from a merchant's own domain indicates a merchant-taken image. For images taken by businesses, those that cannot be automatically identified are marked as having an unknown source. A hash algorithm (pHash) is used to generate fingerprints for the images. A pre-set image theft database is used, which includes fingerprints of historically reported infringing images and fingerprints from blacklists of images stolen by third-party platforms. The generated image fingerprints are compared with the fingerprints in the pre-set image theft database, and the cosine similarity between the fingerprints is calculated. If the similarity exceeds a preset threshold, it is considered a match. If the image fingerprint matches an image in the image theft database, the image source identifier is changed to high risk, and the matching image theft database ID is recorded. The credibility of the data is identified based on the image source and the image theft matching result: official rendered images are considered high credibility (80-100 points), merchant-taken images are considered medium credibility (50-70 points), unknown sources are considered low credibility (20-40 points), and high-risk images are considered the lowest credibility (0-10 points). The collected data is converted into JSON or CSV format to standardize the data format.

[0021] The collected data is preprocessed, and the key fields in the product information table are scanned for null values. If a required field is empty, the record is deleted directly. If a non-required field is empty, the record is retained but marked as a missing field. Keywords are extracted from the product title and attributes, matched with preset category rules, and the matched categories are tagged with category labels.

[0022] S202, based on a large model, performs sentence analysis on the data to generate information credibility scores; Furthermore, Prompt Engineering guides the ChatGLM model to transform the messy raw data into a well-structured and semantically clear standardized description. A standardized rule base is built, defining the size, color, and material of the standard vocabulary. Prompt Engineering inputs raw data and standardized instructions into the ChatGLM, leveraging its contextual understanding capabilities to automatically match the rule base or generate logically consistent standardized expressions. The ChatGLM performs text extraction and image analysis on the data. For text extraction, key information such as price, brand, and material is extracted from product titles and attributes. For image analysis, Optical Character Recognition (OCR) is used to identify price tags, brand logos, and other text in images. A ResNet image classification model is used to identify visual features such as product type and color. ChatGLM compares the semantic consistency between the text description and the visual content. For example, if the image displays bright red while the title description is dark blue, it is considered low consistency; if the title indicates a high-end gift box and the corresponding image shows exquisite and undamaged packaging, it is considered high consistency. Based on the above conflict level comparison, high, medium, and low consistency indicators are output, along with specific conflict explanations.

[0023] Based on the authority of the data source and the accuracy of historical data, each data point is assigned an initial credibility level. For products with high initial credibility and high consistency, the credibility score can be set to 90-100 points. If the consistency is medium, the credibility score drops to 70-80 points. If there is low consistency, even if the initial credibility is high, the final score will drop below 50 points.

[0024] S203. Calculate attribute similarity and semantic similarity based on the standardized data, calculate basic similarity based on attribute similarity and semantic similarity, and correct the basic similarity based on the information credibility score to obtain the final similarity. Specifically, the standardized attribute information is converted into attribute vectors. Categorical attributes (brands) are generated into binary vectors using one-hot encoding; for example, brand A is encoded as [1,0,0] and brand B as [0,1,0]. Numerical attributes (specifications) are mapped to the [0,1] interval through normalization (Min-Max normalization); for example, size XL corresponds to 0.8 and size M corresponds to 0.4. Attribute similarity is then calculated based on the attribute vectors of the two products. ,in, Cosine similarity measures the directional similarity between two attribute vectors in space, with a value range of [−1, 1]. In the context of product similarity, the absolute value is taken as [0, 1]. These are attribute vectors for two products, derived from standardized structured attributes (such as brand, model, specifications, color, etc.). Let be the magnitude of the vector, representing the length of the vector in multidimensional space. The standardized descriptive text (product title, attribute text) is used to generate text vectors using the TF-IDF algorithm. In the TF-IDF algorithm, TF (Term Frequency) is used to count the frequency of words in the document, and IDF (Inverse Document Frequency) measures the general importance of words, with the formula: Where N is the total number of documents. Let t be the number of documents containing word t. The final text vector is the product of TF and IDF. The formula for calculating the semantic similarity between two text vectors T and S is: ,in, Semantic cosine similarity represents the directional similarity between two semantic vectors in space. These are text vectors for two products, generated from descriptive text (such as product titles and attribute descriptions) using the TF-IDF algorithm. Let be the magnitude of the vector, representing the length of the text vector in multidimensional space. The formula for calculating the basic similarity is: (The formula is missing from the provided text). ,in, Based on similarity, , These are the weights for attribute similarity and semantic similarity, respectively, adjusted according to the business scenario (in e-commerce scenarios, the attribute weight is set to 0.6, and the semantic weight is set to 0.4). =1; The formula for correcting the basic similarity based on the information credibility score is: ,in, For the final similarity, Based on similarity, This is a credibility weighting factor.

[0025] S204, Set up a weight mapping table. Before step S203, the system automatically queries based on product category to obtain the optimal weight combination, sorts the products in descending order based on the final similarity calculation results, and obtains a product recommendation list. The large model is used to generate explanatory reasons for the recommendation results in the similar product recommendation table. When there are insufficient similar products or the products are new, a popularity substitution strategy is used. Furthermore, based on the category characteristics of the goods, the system identifies the weight priority of attribute similarity and text similarity. For electronic products, hardware parameters are key to user decisions, so attribute similarity has a higher weight and text description has a lower weight. For clothing and footwear, users pay more attention to text descriptions such as style and material, so text similarity has a higher weight and attribute similarity has a lower weight. For books, users usually select based on text information such as title, author, and synopsis, so text similarity has a dominant weight and attribute similarity has a lower weight. The system stores the mapping relationship between the above categories and weight combinations. The system automatically queries the dynamic weight mapping table based on the category tag of the target product and obtains the weight combination of the corresponding category from the mapping table. When calculating the basic similarity in step S103, the attribute similarity and text similarity are weighted and summed according to the loaded weights.

[0026] The system sorts candidate products in descending order based on the comprehensive similarity after dynamic weight allocation. It then selects the top-ranked products from the sorted list as recommendations, resulting in a similar product recommendation table. When there are insufficient similar products or the products are new, a popularity substitution strategy is activated. This strategy involves selecting substitute products from the target product's category based on popularity indicators such as sales volume, click-through rate, and number of favorites. Similar products are then mixed with best-selling products in a proportional manner for recommendation. A natural language processing model is used to generate explanatory reasons for each recommendation, such as "same brand, same series," etc.

[0027] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: The hybrid calculation of attribute similarity and semantic similarity improves recommendation accuracy compared to a single algorithm; by introducing information credibility score as a weighting factor, similarity calculation is more accurate, reducing the interference of low-credibility products on recommendation results and improving recommendation quality; the dynamic weight allocation mechanism adjusts weights according to product category characteristics, making the recommendation strategy more precise and flexible; and a recommendation system integrating product information credibility assessment, dynamic weight allocation, and negative feedback loop is constructed to achieve accurate, reliable, and adaptive recommendations for candidate products, thereby improving user trust and purchasing efficiency.

[0028] Example 3: In the recommendation methods of Examples 1 and 2, user-expressed disinterest data is typically treated as simple negative samples to reduce the recommendation weight of similar products. This example mines disinterest data, transforming it from a simple rejection signal into a precise scale defining the boundaries of user interests, thereby achieving more accurate and safer recommendations for products of interest. Figure 3 As shown.

[0029] S301, collect negative feedback signals, construct a set of negative feedback commodity vectors based on the negative feedback signals, calculate negative feedback cluster centers based on the set of negative feedback commodity vectors, calculate the distance from candidate commodity vectors to negative feedback cluster centers, and combine the calculated distances with a function to construct a repulsion vector field; Furthermore, negative feedback signals include explicit and implicit feedback. For explicit feedback signals, user-triggered explicit feedback is collected in real time, such as clicking buttons like "not interested" or "recommendation incorrect." For implicit feedback signals, the implicit feedback signal is when a user quickly closes a recommended product after a short browsing session, indicating that the user has no interest in learning more about the product; or when a product remains in a fixed position (such as at the end) in the recommendation list and is continuously ignored by the user, also indicating that the user is not interested in the product. These implicit feedback signals are captured, and the collected negative feedback signals are labeled with intensity. The label intensity is ranked as explicit feedback > quick closure > continuous ignoring. Explicit feedback is the user's direct expression of clear dislike and has the highest intensity; quick closure indicates that the user is not interested in the product. Low interest indicates moderate intensity; continued neglect indicates the user is not very interested in the product, with the lowest intensity. A large-scale model, ChatGLM, is used to map each negative feedback product to a vector, transforming the product's various features into numerical vectors. ChatGLM is a large-scale language model based on deep learning, giving products specific positions and representations in vector space. For example, a clothing product may have features such as color, style, and material. ChatGLM can transform these features into a multi-dimensional vector, forming a set of negative feedback product vectors for the user. This set contains the representations of all products the user is not interested in in vector space. Each negative feedback vector is assigned a weight, determined by the negative feedback intensity and a time decay factor.

[0030] The negative feedback product vector set was clustered using the DBSCAN algorithm, a density-based clustering algorithm that divides the dataset into multiple clusters based on the density of data points. The algorithm can discover clusters of arbitrary shapes. In negative feedback data processing, users' disliked products have various combinations of features, forming clusters of different shapes. After clustering the negative feedback product vector set using the DBSCAN algorithm, the main aversion interest clusters of users are identified. Each aversion interest cluster represents a type of product characteristic that users dislike. For each identified aversion interest cluster, its center vector is calculated. The center vector represents the average characteristic of all vectors in the cluster, that is, the average performance of a certain type of product characteristic that users dislike. The center vector is calculated by taking the mean of all vectors in the cluster. The weighted distance from the candidate product vector to the center of each aversion interest cluster is calculated using Euclidean distance. Then, the distance is adjusted according to the weight of the negative feedback vector. A repulsion function is constructed based on the weighted distance. The value of the repulsion function is negatively correlated with the weighted distance from the candidate product vector to the center of all aversion interest clusters, that is, the closer to the aversion center, the greater the repulsion. The formula for calculating the repulsion function is: ,in, The repulsion force of candidate products, For candidate product vectors, Let i be the center vector of the i-th aversion interest cluster. Let be the weight of the i-th aversion interest cluster, and n be the total number of aversion interest clusters. The Gaussian kernel parameter, σ, controls the range of influence of the repulsion force. The larger the σ, the wider the range of influence of the repulsion force; the smaller the σ, the more concentrated the range of influence of the repulsion force. The repulsion force function transforms the relative positional relationship between candidate products and aversion interest clusters into a quantified repulsion force value, thereby constructing a continuous repulsion field. In the repulsion field, the closer a candidate product is to the center of the aversion interest cluster, the greater the repulsion force it experiences, and the less likely it is to be recommended to the user.

[0031] S302, calculate the minimum distance based on the center of the aversion interest cluster, set a safe distance threshold based on the minimum distance, delineate the interest safety zone in reverse by the center of the aversion interest cluster and the safe distance threshold, and correct the multi-layer interest boundary based on the interest safety zone; Specifically, based on the collected sets of positive and negative feedback items and the calculated centers of aversion interest clusters, the minimum distance from the positive feedback item vector to all aversion cluster centers is calculated: ,in, The minimum distance from the positive feedback product vector to the center of all aversion clusters. For positive feedback product vectors, For the center vector of the j-th aversion interest cluster, a safe distance threshold is set based on the minimum distance from the positive feedback product vector to the center of all aversion clusters. , To establish a safe distance threshold, starting from the user's negative feedback behavior, an interest safety zone is defined in reverse based on the safe distance threshold. This ensures that the distance between products within the interest safety zone and all known aversion interest clusters exceeds the safe distance threshold. The interest safety zone is represented as follows: in, For the product vector, Let i be the center vector of the i-th aversion interest cluster. The safe distance threshold is used as the basis for calculation. The outer boundary of the exploration area calculated in Example 2 is intersected with the safe area defined by the safe distance threshold. The outer boundary of the exploration area in Example 2 is represented as follows: ,in, The radius of the core area, To explore the outer diameter of the area, The effective exploration area is obtained by intersecting the outer boundary of the exploration area with the safe area defined by the safe distance threshold, with the core interest cluster as the center. .

[0032] A concrete example is as follows: Positive feedback product set (products the user likes or has purchased): V1=[120,1] (price 120 yuan, category 1: electronics); V2=[80,3] (price 80 yuan, category 3: daily necessities); V3=[200,2] (price 200 yuan, category 2: clothing); Negative feedback product set (products the user explicitly dislikes): N1=[50,1] (cheap electronics); N2=[30,3] (inferior daily necessities); N3=[1500,4] (expensive luxury goods). Clustering yields the centers of the aversion interest clusters. Using K-Means clustering (assuming k=2), the negative feedback products are divided into two categories: Aversion Cluster 1 (cheap, low-quality products): containing N1=[50,1] and N2=[30,3], with the center: Cneg1= =[40,2];Aversion cluster 2 (high-priced luxury goods): contains N3=[1500,4], center: Cneg2=[1500,4] (single sample cluster), for each positive feedback product Vk, calculate its Euclidean distance to the centers of the two aversion clusters and take the minimum value: (1) Product V1 =[120,1], distance to Cneg1=[40,2]: The distance to Cneg 2 = [1500, 4]: Minimum distance: min(80.00,1393.0)=80.00; (2) Distance from product V2=[80,3] to Cneg1=[40,2]: The distance to Cneg2=[1500,4]: Minimum distance: min(40.01,1420.0)=40.01; (3) Distance from commodity V3=[200,2] to Cneg1=[40,2]: The distance to Cneg2=[1500,4]: Minimum distance: min(160.0, 1300.0) = 160.0, Minimum distance list for positive feedback products: [80.00, 40.01, 160.0], Average: ≈93.34, Median: 80.00, 95th percentile (assuming only 3 samples, take the maximum value): 160.0, Set safe distance threshold: Dsafe=100; Interest safe zone is defined as: That is, the minimum distance from product V to the centers of the two aversion clusters must be greater than 100. To verify whether the positive feedback product is within the safe zone, the minimum distance for V1 is 80.00 (not satisfied), the minimum distance for V2 is 40.01 (not satisfied), and the minimum distance for V3 is 160.0 (satisfied). Only product V3 is within the safe zone. Extended verification: new product V4=[300,1] to Cneg1=[40,2]: To Cneg2=[1500,4]: Minimum distance: 260.0 (satisfying >100), belonging to the safe zone; Modify the multi-layer interest boundary, assuming the original multi-layer interest structure is: the core area is centered on the user's historical interest center Ccore=[100,1.5] with a radius dcore=50, the exploration area: outer diameter... =150, Exploration Zone Boundary: An effective exploration zone must simultaneously meet the following requirements: , Taking V4=[300,1] as an example: Distance to the core area: (If the exploration area does not meet the requirement of ≤150, it was not originally part of the exploration area), adjust the outer diameter of the exploration area to... =200, recalculate the intersection: valid exploration area: Verify that V4 satisfies the exploration zone (200.0 is within [50,200]) and the safe zone (260.0 > 100), and V4 belongs to the valid exploration zone.

[0033] S303, based on the repulsion vector field, corrects the exploration gain coefficient and uses the corrected exploration gain coefficient to generate a safe recommendation list; Furthermore, based on the original exploration gain coefficient, a repulsion force is introduced as a penalty term to reduce the ranking weight of candidate products similar to those disliked by the user, thus avoiding the recommendation of products that the user explicitly dislikes. The revised exploration gain coefficient is: ,in, This is the corrected exploration gain coefficient. These are the original exploration gain coefficients, including the exploration gain coefficients for the core area, exploration area, and outer area. Candidate Products The repulsive force experienced (calculated in step S301) represents the degree of similarity between the product and the one disliked by the user. The repulsion sensitivity coefficient (λ≥0) controls the strength of the repulsion force's influence on the final score. The larger λ is, the stronger the penalty of the repulsion force on the score (suitable for users who are sensitive to negative feedback). The smaller λ is, the weaker the influence of the repulsion force (suitable for users who have a higher tolerance for negative feedback). Based on the modified exploration gain coefficient, candidate products are sorted from high to low, and the top few products are selected to generate a safe recommendation list.

[0034] The technical solutions in the above embodiments of this application have at least the following technical effects or advantages: by using the negative feedback boundary correction mechanism, the characteristics of products that users clearly dislike are actively avoided, reducing recommendation errors. By using dislike to define the boundary of liking, the understanding of the user interest model is more refined and comprehensive. The negative feedback data is transformed from a passive blacklist tool into a precise probe that actively defines the user preference space, thereby achieving more accurate recommendations for products of interest.

[0035] Example 4: Example 1 generates an interest recommendation list by adjusting the recommendation order through exploration gain coefficients. However, interests may shift rapidly due to trending events or sudden demands (such as the surge in demand for hygiene products during a pandemic), leading to a fragmented experience and delayed recommendations. This example, based on the interest recommendation list, generates an intelligent recommendation list by collaboratively using boundary chains and recommendation chains, achieving real-time recommendations and consistent recommendation paths. Figure 4 As shown.

[0036] S401 collects external event data, monitors the intensity of external events, calculates the interest drift index based on historical interaction data, generates boundary chain nodes based on the intensity of external events and the interest drift index, and constructs the boundary chain; Specifically, external event data is collected, including market hotspots, seasonal changes, and breaking news. The collected external data is cleaned to remove noise and invalid data. A real-time monitoring system is established to continuously collect data related to external events and calculate the intensity of these events in real time. User historical interaction data is divided into time periods. For each time period, the number of different categories of user interactions is counted. Based on the proportion of each category in the total number of interactions, the interaction probability of each category is calculated. Information entropy is then calculated based on the historical interaction probabilities, using the following formula: ,in, Historical interaction entropy reflects the diversity and uncertainty of user interaction behavior over a relatively long period of time. For interaction probability, To determine the total number of categories, select user interaction data from a recent period (e.g., the past week or month), calculate the recent interaction entropy using the formula above, and then calculate the interest drift index based on the historical and recent interaction entropies. The formula is as follows: ,in, Interest drift index, For historical interaction entropy, For recent interaction entropy, intensity and drift thresholds are set. When the intensity of an external event exceeds the preset intensity threshold or the interest drift index exceeds the preset drift threshold, a boundary chain node is generated. Each boundary chain node embeds a recommendation chain anchor point. The anchor point contains a path start vector and a chain identifier. The path start vector can be initialized based on the current user state, external event characteristics, etc., and is used to represent the starting direction of the recommendation chain. The chain identifier is used to uniquely identify the boundary chain node. The core area radius is adjusted based on the external event intensity, using the following formula: ,in, The adjusted core area radius, The original core area radius, This is the expansion factor, with a value range of 0 < β < 1. The exploration gain coefficient is adjusted based on the intensity of external events, using the following formula: ,in, This is the corrected exploration gain coefficient. This represents the original exploration gain coefficient. As a correction factor, To determine the intensity of external events, the boundary chain weights are obtained based on event freshness. The boundary chain nodes, the adjusted core area radius, the corrected exploration gain coefficient, and the calculated boundary chain weights are integrated and connected sequentially to construct a complete boundary chain.

[0037] S402, use a prediction model to generate interest nodes, and build a recommendation chain based on the path vectors of the boundary chain nodes and the interest nodes; Furthermore, a graph neural network model is used as the prediction model. Path start vectors are extracted from the boundary chain nodes and input into the prediction model. The prediction model generates interest nodes according to a built-in algorithm. These interest nodes form an interest evolution sequence. Based on the relevance between the interest evolution sequence and the products, the similarity between the products and each interest point in the interest evolution sequence is calculated. Products with high similarity are selected to form a product set. The difference between the feature vectors of two interest points generates a path vector, which represents the path direction. The transition probability from one recommendation node to another, i.e., the transition weight, is calculated. Using the path start vector in the boundary chain node as the starting point of the recommendation chain, an empty recommendation chain structure is created. Based on the generated interest evolution sequence and the transition weights of the recommendation nodes, recommendation nodes are added to the recommendation chain sequentially. Starting from the starting point, the next most likely recommendation node is selected according to the magnitude of the transition weight and connected to the end of the current recommendation chain. This process is repeated until the interest evolution sequence ends, thus constructing the recommendation chain.

[0038] S403, the boundary chain and the recommendation chain are controlled collaboratively, a recommendation score is calculated, an intelligent recommendation list is generated based on the recommendation score, and the boundary chain and the recommendation chain are cross-adjusted. The cross-adjustment is that the boundary chain adjusts the recommendation chain, and the path strength of the recommendation chain adjusts the core area radius of the boundary chain.

[0039] Specifically, the controller monitors the freshness of the boundary chain nodes in real time. This freshness is monitored by continuously tracking the time information of events associated with each node on the boundary chain. The system also focuses on the progress of each recommendation node on the recommendation chain, counting the number of completed (i.e., viewed or selected by the user) recommendation nodes and the number of remaining incomplete recommendation nodes to determine the completion rate of the recommendation chain path. The boundary chain weight is multiplied by the boundary chain position score, and the recommendation chain weight is multiplied by the recommendation chain path score. These two products are then added together to obtain a recommendation score that comprehensively considers both boundary chain and recommendation chain factors. Based on this recommendation score, a smart recommendation table is generated. The recommendation score measures the suitability of the recommended products or content for the user; a higher score indicates that the recommendation better matches the user's current interests. Requirements: The controller monitors the status of boundary chain nodes in real time. When a boundary chain node switch is detected, such as when a user's interest shifts from one type of product to another, or when a hot event causes the boundary chain node to update, the controller forces the recommendation chain to reinitialize. This means abandoning the currently advancing recommendation chain path and regenerating the interest evolution sequence and recommendation nodes based on the user's current interests represented by the new boundary chain node. After the recommendation chain is displayed to the user, user feedback on their recommended path selection is collected. Based on this feedback, the path strength of the recommendation chain is comprehensively evaluated. Path strength reflects the user's acceptance and approval of the recommended path. The formula for adjusting the core area radius of the boundary chain based on path strength is: ,in, To use the core area radius adjusted by path strength, The adjusted core area radius, These are pre-set adjustment parameters used to control the degree of influence of path intensity on the core area radius. This represents path strength.

[0040] The technical solutions described in the above embodiments of this application have at least the following technical effects or advantages: In emergency scenarios, dual-chain collaboration makes the recommendation score more accurate, and the boundary chain and collaborative controller can adjust the interest boundary and recommendation path in real time, providing a more humanized, coherent and accurate e-commerce recommendation solution, moving from static division to dynamic guidance.

[0041] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A product recommendation method based on big data, characterized in that, include: S101 collects users' historical interaction behavior data, identifies the core area based on the historical interaction behavior data, delineates the exploration area based on the core area, identifies the peripheral area based on the exploration area, and constructs a multi-layered interest boundary based on the core area, exploration area, and peripheral area; S102, collect product data, calculate attribute similarity and semantic similarity based on product data, calculate basic similarity through attribute similarity and semantic similarity, correct the basic similarity to obtain the final similarity, obtain the basic fit layer score based on the final similarity, and calculate the exploration gain coefficient layer by layer based on the multi-layer interest boundary. S103 generates an interest recommendation list based on the exploration gain coefficient, tracks user behavior data in real time, and updates the multi-layer interest boundaries.

2. The product recommendation method based on big data as described in claim 1, characterized in that, The steps for calculating the final similarity are as follows: convert the standardized attribute information into attribute vectors and calculate attribute similarity; generate text vectors from the standardized description text using an algorithm and calculate semantic similarity; calculate the basic similarity based on the weighted combination of attribute similarity and semantic similarity, combined with the information credibility score; and correct the basic similarity to obtain the final similarity.

3. The product recommendation method based on big data as described in claim 2, characterized in that, The formula for calculating attribute cosine similarity is: ,in, Cosine similarity represents the directional similarity between two attribute vectors in space. Let be the attribute vectors of two products. Let be the magnitude of the vector, representing the length of the vector in the multidimensional space; the formula for calculating semantic similarity is: ,in, Semantic cosine similarity represents the directional similarity between two text vectors in space. Given the text vectors of two products, Let be the magnitude of the vector, representing the length of the text vector in multidimensional space. The formula for calculating the basic similarity is: (The formula is missing from the provided text). ,in, Based on similarity, , These are the weights for attribute similarity and semantic similarity, respectively. =1; The formula for correcting the basic similarity based on the information credibility score is: ,in, For the final similarity, Based on similarity, This is a credibility weighting factor.

4. The product recommendation method based on big data as described in claim 1, characterized in that, The steps in S101, which involve identifying the core area based on historical interaction data, defining the exploration area based on the core area, and identifying the outer area based on the exploration area, are as follows: Identify the core area, exploration area, and outer area based on the calculated Euclidean distance. If d ≤ d represents the Euclidean distance, and the core area is centered on the user's core interest point. A spherical region with radius ; if d≤ If the item falls into the exploration zone, the exploration zone is defined by its inner diameter being the radius of the core area. The outer diameter is Annular region; If d If the item falls into the outer zone, the outer zone is the area outside the exploration zone. All vector spaces other than that.

5. The product recommendation method based on big data as described in claim 1, characterized in that, The method for constructing multi-layer interest boundaries in S101 is as follows: by identifying high-frequency user interaction behaviors, core interest points are calculated, and core radius is defined by the core interest points; the core radius is used as the inner diameter of the exploration area, and the outer diameter of the exploration area is obtained through the core interest points, thereby defining the exploration area; all vector spaces outside the outer diameter of the exploration area are defined as the outer area.

6. The product recommendation method based on big data as described in claim 1, characterized in that, The method for calculating the exploration gain coefficient in step S102 is as follows: The formula for calculating the exploration gain coefficient is: ,in, For the exploration gain coefficient of the core area, Based on the basic fit layer score, As a core area reward factor; ,in, This represents the exploration gain coefficient for the exploration zone. Based on the basic fit layer score, Scoring of the boundary extension layer To explore the incentive coefficient To explore the attenuation factor; ,in, For the exploration gain coefficient of the outer region, Based on the basic fit layer score, It is a very small constant.

7. The product recommendation method based on big data as described in claim 1, characterized in that, The recommended list is generated based on the exploration gain coefficient and also includes: S301, collect negative feedback signals, construct a set of negative feedback commodity vectors based on the negative feedback signals, calculate negative feedback cluster centers based on the set of negative feedback commodity vectors, calculate the distance from candidate commodity vectors to negative feedback cluster centers, and combine the calculated distances with a function to construct a repulsion vector field; S302, calculate the minimum distance based on the center of the aversion interest cluster, set a safe distance threshold based on the minimum distance, delineate the interest safety zone in reverse by the center of the aversion interest cluster and the safe distance threshold, and correct the multi-layer interest boundary based on the interest safety zone; S303, based on the repulsion vector field, corrects the exploration gain coefficient and uses the corrected exploration gain coefficient to generate a safe recommendation list.

8. The product recommendation method based on big data as described in claim 7, characterized in that, The step of collecting negative feedback signals involves collecting explicit feedback and implicit feedback, with the explicit feedback having a higher strength than the implicit feedback.

9. The product recommendation method based on big data as described in claim 7, characterized in that, The method for constructing the repulsive vector field is as follows: The formula for calculating the repulsive force function is: ,in, The repulsion force of candidate products, For candidate product vectors, Let i be the center vector of the i-th aversion interest cluster. Let be the weight of the i-th aversion interest cluster, and n be the total number of aversion interest clusters. The Gaussian kernel parameter is used to control the range of influence of the repulsive force. The relative positional relationship between the candidate product and the aversion interest cluster is transformed into a quantified repulsive force value through the repulsive force function, thereby constructing a continuous repulsive vector field.

10. The product recommendation method based on big data as described in claim 1, characterized in that, The method further includes: S401 collects external event data, monitors the intensity of external events, calculates the interest drift index based on historical interaction data, generates boundary chain nodes based on the intensity of external events and the interest drift index, and constructs the boundary chain; S402, use a prediction model to generate interest nodes, and build a recommendation chain based on the path vectors of the boundary chain nodes and the interest nodes; S403, the boundary chain and the recommendation chain are controlled collaboratively, a recommendation score is calculated, an intelligent recommendation list is generated based on the recommendation score, and the boundary chain and the recommendation chain are cross-adjusted. The cross-adjustment is that the boundary chain adjusts the recommendation chain, and the path strength of the recommendation chain adjusts the core area radius of the boundary chain.