Product decision optimization method and device based on mapping knowledge domain, equipment and medium
By building a knowledge graph and user feature portraits, and dynamically updating the decision model based on feedback information, the problem of the recommendation system being unable to update in real time in existing technologies is solved, and personalized and adaptive product recommendations are achieved.
Patent Information
- Application Number
- CN202510871410.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-26
AI Technical Summary
Existing recommendation technologies are unable to update knowledge graphs and decision models in real time, resulting in the inability of recommendation systems in insurance, finance, and healthcare scenarios to dynamically match user behavior and product structure, making it difficult to achieve personalized recommendations.
By acquiring user, product and environment information, building a knowledge graph, generating user feature portraits, and training decision models based on preset incentive mechanisms, feedback information is collected to dynamically update the model and graph.
It improves the pertinence of recommendation results and the system's adaptability, and achieves real-time response and personalized matching to user preferences.
Smart Images

Figure CN120707188A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a product decision optimization method, device, equipment and storage medium based on knowledge graph. Background Art
[0002] In the field of fintech business, the insurance industry has developed rapidly in recent years, with increasingly diverse product forms and expanding customer coverage. However, the current mainstream insurance product recommendation systems generally rely on static rules and simple user tags, which makes it difficult to meet the diverse and dynamic protection needs of customers. Especially in scenarios facing complex populations, such as newlyweds, high-net-worth customers, or users holding multiple insurance policies, traditional recommendation methods are mainly based on static information (such as age, gender, and insurance history) for product matching. They are unable to dynamically capture changes in customer behavior or personalized preferences, resulting in repeated and redundant recommendations and a poor customer experience. At the same time, the insurance products themselves have complex terms, multiple protection dimensions, and long coverage periods, which are difficult for customers to accurately understand. It is also difficult for the system to complete accurate reasoning and differentiated recommendations based on product knowledge.
[0003] In the medical and health business, with the improvement of residents' health awareness and the increase in demand for personalized health management, the demand for health insurance products has grown rapidly, but existing recommendation technologies have difficulty in effectively linking the automatic matching of customer health status and protection plans. On the one hand, the system lacks the ability to deeply structure and dynamically model health data (such as physical examination reports, medical records, chronic disease medication status, etc.), resulting in inaccurate individual health risk assessments; on the other hand, it is difficult to dynamically identify the behavior patterns and protection preferences of customers with different health characteristics, and the recommendation results are likely to deviate from actual needs. In addition, in the face of new products, new policies or clinical standard updates, the recommendation model cannot respond in a timely manner, affecting business response efficiency and product penetration.
[0004] In the general use of intelligent recommendation systems, existing technologies have the following major shortcomings: First, the granularity of user profile construction is coarse, and it is impossible to integrate behavioral, attribute, and relationship information from multi-source heterogeneous data, resulting in insufficient state expression capabilities; second, most recommendation models are based on static optimization and lack continuous learning and dynamic update mechanisms, making it difficult to adapt to evolving customer interests and changes in the external environment; third, recommendation decisions rely on limited rules or shallow features, and cannot use knowledge reasoning to determine the deep logical matching relationship between user needs and products. In addition, although structured semantic networks such as knowledge graphs have been introduced in some recommendation tasks, the update mechanism is still mainly based on periodic manual updates, which makes it difficult to reflect users' latest behaviors in real time, limiting the system's ability to accurately respond to user preferences.
[0005] In summary, existing recommendation technologies still have significant deficiencies in modeling depth, feedback response, knowledge expression and update mechanisms, making it difficult to achieve dynamic recommendations for complex user behaviors and product structures. The intelligent matching capabilities in insurance, finance, and healthcare scenarios need to be improved urgently. Summary of the Invention
[0006] The main purpose of the present invention is to provide a product decision optimization method, device, equipment and storage medium based on knowledge graph, aiming to solve the technical problem that the existing technology cannot use user behavior feedback in real time to update the knowledge graph and decision model, resulting in the recommendation system being unable to form a closed-loop optimization and it is difficult to continuously improve the accuracy and personalization of recommendations.
[0007] To achieve the above objectives, the present invention provides a product decision optimization method based on knowledge graph, comprising:
[0008] Obtain and integrate user-related information, product-related information, and environmental information to generate target information;
[0009] Constructing a knowledge graph based on the target information;
[0010] Constructing a user feature profile based on the user-related information and the knowledge graph;
[0011] The user feature profile and the information in the knowledge graph are used as input states, the product information set is used as the action space, and a decision model is trained and generated based on a preset incentive mechanism;
[0012] outputting a target project through the decision model;
[0013] Collect feedback information on the target project, and update the knowledge graph and the decision model based on the feedback information.
[0014] Furthermore, to achieve the above objectives, the present invention provides a product decision optimization device based on a knowledge graph, comprising:
[0015] Target information generation module, used to obtain and integrate user-related information, product-related information and environmental information to generate target information;
[0016] A graph construction module, used to construct a knowledge graph based on the target information;
[0017] A user portrait construction module is used to construct a user feature portrait based on the user-related information and the knowledge graph;
[0018] A decision model training module is used to train and generate a decision model based on a preset incentive mechanism, taking the user feature profile and the information in the knowledge graph as input states and the product information set as an action space;
[0019] A target item reasoning module, configured to output a target item through the decision model;
[0020] A model updating module is used to collect feedback information on the target project and update the knowledge graph and the decision model based on the feedback information.
[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a knowledge graph-based product decision optimization program stored in the memory and run on the processor. When the knowledge graph-based product decision optimization program is executed by the processor, the steps of the knowledge graph-based product decision optimization method as described above are implemented.
[0022] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a product decision optimization program based on a knowledge graph is stored. When the product decision optimization program based on a knowledge graph is executed by a processor, the steps of the product decision optimization method based on the knowledge graph as described above are implemented.
[0023] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a product decision optimization method, device, equipment and medium based on knowledge graph, including: acquiring and integrating user-related information, product-related information and environmental information to generate target information; constructing a knowledge graph based on the target information; constructing a user feature portrait based on the user-related information and the knowledge graph; using the user feature portrait and the information in the knowledge graph as input states, and the product information set as the action space, to train and generate a decision model based on a preset incentive mechanism; outputting the target project through the decision model; collecting feedback information on the target project, and updating the knowledge graph and decision model based on the feedback information. The present invention improves the comprehensiveness of the input state through a fusion modeling method that includes user information, product information and environmental information, and constructs a dynamic update mechanism in combination with the knowledge graph and user feedback to drive the continuous optimization of the decision model, thereby improving the pertinence of the recommendation results and the adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0025] Figure 1 A schematic diagram of an application environment of a product decision optimization method based on a knowledge graph in an embodiment of the present invention;
[0026] Figure 2 This is a flow chart of an embodiment of a product decision optimization method based on knowledge graph according to the present invention;
[0027] Figure 3 This is a functional module diagram of a preferred embodiment of the product decision optimization device based on knowledge graph of the present invention;
[0028] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0029] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0031] The product decision optimization method based on knowledge graph provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain and integrate user-related information, product-related information and environmental information through the user terminal to generate target information; construct a knowledge graph based on the target information; construct a user feature profile based on the user-related information and the knowledge graph; use the user feature profile and the information in the knowledge graph as input states, use the product information set as the action space, and train and generate a decision model based on a preset incentive mechanism; output the target project through the decision model; collect feedback information on the target project, and update the knowledge graph and decision model based on the feedback information. The present invention improves the comprehensiveness of the input state through a fusion modeling method that includes user information, product information and environmental information, combines the knowledge graph and user feedback to build a dynamic update mechanism, drives the decision model to be continuously optimized, and thereby improves the pertinence of the recommendation results and the adaptability of the system. Among them, the user terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0032] See also Figure 2 , Figure 2 This is a flowchart of an embodiment of the product decision optimization method based on knowledge graph provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than here.
[0033] like Figure 2 As shown, the product decision optimization method based on knowledge graph proposed in the present invention includes the following steps:
[0034] S10, acquiring and integrating user-related information, product-related information, and environmental information to generate target information;
[0035] In this embodiment, user-related information refers to a data set directly associated with an individual user, including the user's basic identity attributes, historical behavior records, interaction preferences, and historical purchase behavior. Basic identity attributes include, but are not limited to, age, gender, region, occupation, marital status, and family structure, and can be sourced from user registration information or data provided by a real-name authentication platform. Behavior records include page views, clicks, dwell time, search keywords, and actions such as adding items to a shopping cart. Interaction preferences can be dynamically generated using user behavior pattern mining algorithms, such as interest distribution based on click frequency. Purchase behavior includes transaction data on financial or medical platforms, such as insurance registration records, product renewals, and policy cancellations.
[0036] Product-related information refers to the set of attributes for the services or products managed by the system, consisting of structured product attribute fields and unstructured textual descriptions. Attribute fields such as product type, price range, coverage, premium structure, compensation ratio, eligible population, and policy period can be sourced from product design databases or industry-standard coding systems. Unstructured descriptions, such as terms and conditions, usage restrictions, and sales pitches, can be structured by extracting keywords and semantic vectors through natural language processing.
[0037] Environmental information refers to the peripheral contextual content that influences user decisions or behaviors, encompassing temporal, geographical, social, and policy environments. Temporal context, which can include labels like holidays, weekdays, and seasons, influences user attention to coverage periods and short-term products. Geographical context, such as the user's current location, climate in their home city, and the density of medical resources, is particularly important in health insurance recommendations. Social context refers to the characteristics of the user's community, the media content they follow, and their interactions within their social circle, which can be extracted through multi-source information aggregation. The policy context involves current regulatory guidelines for financial or medical services, legal changes, and updates to industry standards, and can be used as a constraint in generating target information.
[0038] Multi-source information is obtained from various distributed databases or real-time collection channels through the data synchronization module. After acquisition, the data needs to be preprocessed to adapt to a unified processing flow. Format standardization operations include timestamp format conversion, code mapping normalization, and unit conversion to ensure structural consistency across data. The data cleansing process uses a rules engine and anomaly detection algorithms to identify missing fields, format errors, and extreme value records, and repairs them through methods such as filling, elimination, and interpolation. Structural transformation uses ETL processes and semantic modeling to uniformly map data into a standard set of fields. It also supports the modeling of primary key or foreign key relationships between fields to form a standardized data table structure.
[0039] Target information is generated based on the standardized data tables obtained through the aforementioned preprocessing. Field selection, feature construction, and relationship extraction are used to create a unified expression vector for downstream use. Field selection filters unnecessary information based on its relevance to the recommendation task, retaining factors that significantly influence recommendation decisions. Feature construction includes operations such as single-field normalization, cross-field combination to generate cross-features, and multidimensional feature tensor generation. Relationship extraction, based on knowledge modeling strategies, constructs candidate association paths between users and products, providing a structural basis for subsequent knowledge graph construction.
[0040] By building a data collection platform, users, product data, and environmental data can be synchronously acquired from various data sources. This data can be localized and standardized by combining a Kafka-based real-time stream processing engine with a Hive-based data warehouse system. Format unification can be achieved by introducing a universal schema parser and a multi-language format conversion engine, mapping heterogeneous sources such as JSON, CSV, and XML to an internal standard model. Missing value handling can be accomplished through a context-dependent infill model. For example, missing data in the age field can be predicted and filled based on the occupational and regional distribution of similar populations. Outlier detection can be achieved by combining box plot analysis with a clustering-based outlier identification mechanism. Structural transformations support multi-table concatenation and field naming standardization through Spark SQL and pre-set field mappings, while maintaining associated primary indexes based on business primary keys. Target information generation can be centrally accomplished through the construction of a feature engineering module, including standardized encoding of static fields, sliding window statistics and frequency modeling for dynamic behavior fields, TF-IDF weight generation for text fields, and feature importance assessment for comprehensive scoring models.
[0041] By introducing a data tag management system, user behavior data and product data can be semantically labeled uniformly, enabling controlled cross-domain data integration. The system can pre-set a tagging system and dynamically adjust tag granularity and coverage based on domain requirements. The tag-based feature construction process facilitates the establishment of explicit or implicit connections between users and products. A graph-structured intermediate layer can also be introduced to construct an initial user-feature-product triple network during the target information generation phase, forming a candidate relationship pool to support the initialization of subsequent graph computation modules.
[0042] Target information is a unified data representation used to support user recommendation decision-making. Essentially, it is a set of high-quality semantic data that integrates three core categories of information: user, product, and environment. This information is preprocessed, standardized, cleansed, and structured to create a high-quality semantic data carrier. The construction of target information aims to break down information silos and enable heterogeneous data to collaboratively participate in the derivation of recommendation strategies within a unified logical model.
[0043] Taking the medical and health scenario as an example, suppose a user has recently frequently browsed information related to the management of hypertension and chronic diseases, and recently registered with a cardiologist. The health record records that his BMI is high, and the user's residence is in the high temperature warning stage. In addition, the user has previously purchased physical examination services but has not yet configured health insurance. Based on the above behavioral trajectories and historical transaction records, the platform aggregates and normalizes its user attributes (age, gender, health risk labels), interactive behaviors (page browsing keywords, recent registration departments), service records (physical examination purchase records), product data (currently sold chronic disease insurance, health management services coverage, premiums, protection time limit, etc.) and environmental information (weather conditions, regional medical resource accessibility, public health warning index). Through time synchronization and structured mapping, these data are uniformly encapsulated into a target information structure.
[0044] In this structure, the user field records individual ID, health risk factors, multi-dimensional tags, and summaries of recent interactions. The product field includes structured summary information for multiple candidate health products, such as covered diseases, waiting periods, service availability, and historical insurance coverage rates. The environment field records the current geographic location's medical coverage density, seasonal risk levels, and the popularity of health topics similar users follow on social media. These fields undergo field cleaning, missing information completion, and unified unit conversion, and are encapsulated using feature vectors to form the input foundation for the recommendation system. This target information is not only descriptive but also has strong upstream and downstream linkage capabilities. It can be read by subsequent knowledge modeling modules and used to construct user profiles and calculate recommendation strategies.
[0045] This implementation integrates user data, product data, and environmental data, and constructs unified target information for intelligent recommendations through standardization, cleansing, and structural transformation. This not only improves the accuracy and reusability of information processing, but also provides a high-quality foundational data representation for downstream knowledge modeling and personalized recommendations. This operational process effectively addresses issues such as information redundancy, inconsistent formats, and disconnected fields, improving data fusion efficiency and semantic consistency.
[0046] S20, constructing a knowledge graph based on the target information;
[0047] In this embodiment, the process of building a knowledge graph relies on previously generated target information, which includes structured data fields such as user information, product information, and environmental information. To achieve semantic expression and relationship modeling of information, the entities, attributes, and interactions in the target information need to be parsed and organized layer by layer.
[0048] First, identifying entities from the target information is the foundation of this process. Entities can include users, products, service items, disease tags, insurance items, and operational behaviors. These entities typically have unique identifiers and structural attributes. For example, a user entity might include user ID, risk preferences, and behavioral tags; a product entity might include the insurance name, applicable population, coverage, and premium calculation rules.
[0049] After identifying an entity, its structural attributes must be extracted as its label attributes. Entity attributes exist as fields in the target information, such as "user age," "product covered diseases," and "product waiting period." The attribute set for each entity must maintain field consistency and type validity, and any missing or abnormal fields must be verified and repaired.
[0050] Semantic relationships between entities can be extracted from user-product interaction records. For example, a user's actions such as "browsing," "inquiring," "purchasing," and "rejecting" a product can create behavioral edges between the user and product entities. Another example is the relationship edges between different products due to their shared service to the same demographic or overlapping product coverage. These relationship edges are typically generated using rule engines or graph-building strategies, drawing on multiple dimensions of the target information, including behavioral logs, transaction records, and product structure descriptions.
[0051] Finally, the aforementioned entities are represented as nodes in a graph, and the relationships as edges, creating a structured representation. The entire graph is then written into a graph database for storage and management. The graph database uses a graph-oriented storage engine (such as Neo4j, ArangoDB, or an RDF-based triple structure) to support subsequent graph queries, graph traversals, and embedded computational tasks.
[0052] In actual construction, in order to improve update efficiency and graph timeliness, when new data is received in the target information, the incremental update process can be triggered on demand to automatically detect the new entities and their conflicting relationships with the existing graph, and only update the new nodes or changed edges to avoid repeated reconstruction of the entire graph.
[0053] In practical applications, knowledge graphs can be constructed and updated in a variety of ways across different system scenarios. For example, in financial business scenarios, an entity extraction and synchronization strategy based on the ETL (Extract-Transform-Load) process can be adopted. Customer information, transaction records, and product catalog fields in the core business database can be extracted and mapped into a graph structure. The relationship rule module then automatically generates edges between customers and products, and introduces preference strength as a weight attribute written into the edge object.
[0054] It is also possible to use a graph to build a scheduling module in scenarios where heterogeneous data sources exist. The graph can be constructed in stages into user subgraphs, product subgraphs, and behavior subgraphs, and then merged into a complete graph to achieve distributed storage and parallel updates.
[0055] Example: In the healthcare domain, the target information for patient Xiao Zhang includes their recent appointment with the endocrinology department, their physical examination report indicating elevated blood sugar, and their search history for diabetic diet and medications. The system identifies entities such as "User_Xiao Zhang," "Department_Endocrinology," "Disease_Diabetes," and "Product_Diabetes Special Insurance Plan" from this information, and creates edges such as "Registered at," "Searched for," "Physical Examination Indicator Relevance," and "Recommended Insurance Product." These nodes and edges are then incorporated into the healthcare knowledge graph, forming a structured diagram that can be used for reasoning and recommendation.
[0056] In the fintech business domain, the target information for customer Zhang indicates that he is 28 years old, currently a programmer, has purchased one-year term life insurance, and has browsed pages for pension annuity products on the platform. The system identifies the entities "User_Zhang," "Occupation_Programmer," "Product_Term Life Insurance," and "Product_Annuity Insurance," constructing edges such as "Purchased," "Browsed," and "Suitable Occupation," storing these edges as a financial customer profile graph for exploring future insurance purchase possibilities or developing differentiated operational strategies.
[0057] This implementation automatically identifies entities and attributes from structured target information and establishes semantic edges based on interaction behavior and product structure, creating a graph structure with contextual semantic expression capabilities. This enables the fusion modeling of user behavior data and product knowledge. Compared to traditional relational data or tabular recommendation models, this graph structure possesses stronger upstream and downstream information transmission capabilities and high-level semantic reasoning capabilities, providing high-quality structured support for subsequent user profile generation and intelligent recommendation strategies.
[0058] S30, constructing a user feature profile based on the user-related information and the knowledge graph;
[0059] In this embodiment, user profiles are constructed based on user-related information and information in the knowledge graph, aiming to integrate the user's static attributes and dynamic behaviors. User-related information typically includes the user's registration information on the platform, behavior logs, transaction records, etc. Registration information such as age, gender, and occupation are static attributes, while behavior logs and transaction records reflect the user's preferences, needs, and risk characteristics within a specific time window, forming a dynamic representation.
[0060] When parsing static attribute data from user-related information, it's necessary to define a standardized field structure and construct a unified feature vector representation. Static fields such as "gender" can be represented using binary discrete values, "age" can be converted into age group labels, and "occupation" can be coded based on the occupational classification system. If missing or ambiguous fields are encountered during processing, supplementary values can be generated using methods such as data interpolation, rule inference, or user interaction.
[0061] When integrating knowledge graph information, we first need to locate the entity node in the knowledge graph that corresponds to the current user. Then, we extract the node's direct connections and adjacent entities, such as associated products, diseases, and services, from the graph structure. The information carried by these adjacent nodes (such as product category, interaction behavior type, product weight, disease label, etc.) can be organized into a graph context feature set.
[0062] The static attribute feature set is fused with the graph context feature set. The final fused feature vector can be generated using methods such as concatenation, multi-layer perceptron fusion, and attention-weighted fusion. The fusion process should ensure feature dimension consistency and semantic consistency. If necessary, discrete fields should be converted into continuous representation vectors through an embedding layer.
[0063] The fused feature vectors must be uniformly encoded to form a standardized user profile. Encoding methods may include feature hashing, normalization, one-hot encoding, vector reduction, and other methods, adjusted based on the input structure of the subsequent model. The final user profile should structurally include basic attribute dimensions, behavioral preference dimensions, and graph context dimensions, and be provided to the recommendation module in the form of a tensor or structure.
[0064] In practice, constructing user profiles can employ different technical approaches depending on the system architecture. In financial services systems, the user's basic profile information and insurance policy transaction records from the past three months are first retrieved from the customer management system. Next, the user's recently browsed, insured, and consulted product nodes are extracted from the knowledge graph, along with the associated coverage items and user behavior weights. A structured feature vector is then generated by concatenating the basic fields with the graph fields. This is then encoded using L2 normalization, ultimately forming a standard profile vector for use in the recommendation strategy network.
[0065] In healthcare scenarios, the construction of user profiles relies more heavily on diagnosis and treatment records and external knowledge structures. Entity alignment can be performed between a patient's historical medical records and the medical knowledge graph to obtain corresponding high-frequency disease labels and related service entities. For example, if a patient has frequently visited a cardiology department and taken cardiovascular medications in the past six months, potential disease risk nodes in the graph can be inferred and embedded as features within the graph context. This information, such as health scores and treatment preferences, is then integrated to form a multi-dimensional feature vector, which is then compressed into a uniform-length encoding vector through an embedding mapping network for use in subsequent model modules.
[0066] In multi-source data scenarios, it is also possible to splice and clean user-related information from different data systems, standardize the field structure through the feature alignment layer, and ensure that fields from the CRM system, behavior log system, and graph system can be integrated and calculated.
[0067] This embodiment integrates the user's static attribute information with the contextual association information in the knowledge graph to form a model. The user feature portrait can not only accurately depict the user's individual characteristics, but also reflect their behavioral preferences and semantic relationships, significantly improving the subsequent recommendation system's ability to perceive user needs.
[0068] S40, using the user feature profile and the information in the knowledge graph as input states, and the product information set as an action space, and training and generating a decision model based on a preset incentive mechanism;
[0069] In this embodiment, the user profile and the information in the knowledge graph together constitute the input state, which means the encoded representation of the decision model's perception of the current context. Among them, the user profile contains static attribute fields and graph interaction vectors, while the information in the knowledge graph can be reflected as the position state of the user's current behavior in the graph structure, including the edge features between the user and high-frequency nodes, relationship type weights, path depth statistics, etc. The final representation of the input state can be a concatenated vector, a structural tensor, or an embedded representation obtained after encoding through a graph neural network.
[0070] The product information set is a candidate set of all recommended objects, considered a discrete action space. Each product has a unique identifier in the action space. The product's structural attributes (such as coverage, applicable population, and risk level) can be used as additional features in the network modeling, but this does not change the discrete nature of the action space.
[0071] When building a decision-making model, you need to design a policy network and a value network. The policy network takes a state vector as input and outputs a probability distribution for each action in the product set, indicating the tendency to recommend each product under the current state. The value network takes a state and action pair as input and outputs an estimated reward for each state-action pair, which is used to assess the long-term value of an action under the current state. Both types of networks can use a multi-layer perceptron structure or enhance the expressiveness of state representations based on graph neural networks or attention mechanisms.
[0072] The incentive mechanism maps user feedback into a quantitative score by defining a reward function, which is then used in the training process. The reward function should weight feedback differently based on business needs. For example, purchase behavior could be assigned a high positive reward, browsing behavior a moderate reward, and rejection a negative penalty. Reward values can be discounted over time, creating a cumulative discount reward structure commonly used in reinforcement learning to support long-term behavioral learning.
[0073] The training process uses policy gradient optimization methods, such as REINFORCE or the Actor-Critic architecture, to update network parameters by sampling user interaction trajectories and combining them with a reward function. This process continuously improves the policy network's ability to select high-reward actions under different states, thereby forming a model with adaptive recommendation capabilities.
[0074] In financial business systems, decision models can be built and trained through the following implementation path: first, extract fields such as occupation, income, and previous insurance records from the user portrait, and obtain the interaction path characteristics between the user and the purchased products from the knowledge graph, and splice them to form the input state; define all insurance products provided by the platform as a set of discrete actions; build a two-tower structure, one of which is a state encoding network and the other is an action encoding network, and output the recommendation probability distribution through the inner product mechanism; generate recommendation results through sampling and record user feedback. If insurance is purchased, a reward of +1 is given, and if it is skipped, a value of -0.5 is assigned. The strategy tower parameters are optimized using the REINFORCE method.
[0075] In medical scenarios, disease risk nodes from patient medical records and graphs can be extracted and encoded as states. A collection of hospital services or checkup packages can be used as the action space. A reward function based on service completion rate and patient satisfaction is designed, with higher rewards for service acceptance and satisfactory feedback, and negative rewards for non-selection or withdrawal. The training process can be iteratively updated based on an actor-critic structure, allowing the model to gradually learn to recommend appropriate service combinations for users with different health profiles.
[0076] It is also possible to introduce multimodal information fusion mechanisms in cross-domain scenarios, such as fusing graph structure embedding and semantic embedding information to construct a state vector, adding the BERT vector of the product description text to the product action space to enhance the action representation, and integrating the outputs of multiple models into a unified policy network through policy distillation methods.
[0077] This embodiment combines user profiles with graph information to construct state inputs and defines product information sets as an action space, helping the model capture the complex connections between user behavior context and product semantics. By introducing a feedback-based incentive mechanism and a policy optimization structure, the model continuously updates and adapts to changing user preferences and market dynamics. Compared to traditional static recommendation methods, it achieves higher personalized matching and recommendation accuracy, dynamically improving user experience and interaction conversion efficiency.
[0078] S50, outputting the target project through the decision model;
[0079] In this embodiment, outputting the target item is the process of performing an inference operation after completing the decision model training, combining the current user features with the knowledge graph information. This operation first requires obtaining the latest feature representation of the current user, including static attributes and interaction behavior summaries, and combining it with the real-time structural information in the graph to generate an inference state vector. The inference state vector can be constructed in various ways, such as directly concatenating the user feature vector and the graph embedding vector, or dynamically fusing the contextual information of its associated entities through the attention mechanism to form an encoded representation of the current user state.
[0080] After receiving the inference state vector, the decision model generates a probability distribution defined over the action space at its output layer. Each action corresponds to a candidate item, and its probability value reflects the likelihood of that item being recommended under the current state. If a policy network is used, the output is a Softmax-normalized probability vector. If a value network is used, the state-action value is mapped to a sampling probability using either the Boltzmann strategy or the ε-greedy strategy.
[0081] Determining the target item is typically done through a probabilistic sampling mechanism, where a product is selected from the aforementioned distribution based on probability as the recommended result. A top-K strategy can also be used to re-rank the results before performing sampling. The target item identifier is a unique number for the product in the action space. This is mapped to structured product information by querying the product library, ultimately forming the target item entity that can be presented to the user.
[0082] To ensure the consistency and explainability of reasoning, the output results, in addition to the target item, can also include information such as the current state vector, corresponding action score, path reasoning link, etc., which can be used for subsequent result tracking, strategy analysis and user feedback modeling.
[0083] In a health insurance recommendation system for family customers, target project output can be achieved in the following ways: first, extract the current user's identity label, disease risk category, previous insurance record and other vectors from the user portrait service, and construct an inference state by combining the service nodes with the most recently interacted in the knowledge graph; input this state vector into the trained policy network to obtain the recommendation probability distribution of all insurance products in the current state; select the top several products with the highest scores through probabilistic sampling; extract the complete details of the product from the product information system based on its identification, and finally use the selected product combination as the target project for use by the display module.
[0084] In financial scenarios, such as smart investment advisory platforms, after users adjust their asset allocation preferences, the system extracts characteristics such as their risk tolerance, historical holding performance, and financial knowledge level, and integrates them with market trend nodes and historical trading behavior nodes in the knowledge graph to construct the current state; this state is sent to the reinforcement learning model inference module to generate an action probability distribution including fund portfolios, insurance configurations, and trust recommendations; the system selects the current optimal investment advisory plan based on the strategy output, maps it into a structured investment advice plan, and returns it to the user-side page for display.
[0085] A fine-scale filtering mechanism for candidate sets can also be introduced into the service platform. That is, before probability sampling, the rule engine is used to filter out action items that are incompatible with the user's current status, such as products already held and products with legal restrictions, and the action selection process is only performed on the remaining candidate items to improve the feasibility of recommendations.
[0086] This embodiment generates an inference state by inputting current user characteristics and graph structure information, and outputs the optimal target item in the action space based on the decision model, enabling dynamic response to user personalized needs and contextual scenarios. This method combines a collaborative mechanism of graph structure perception, behavior-driven reasoning, and model strategy optimization, significantly improving recommendation accuracy, adaptability, and system responsiveness compared to traditional recommendation methods based on static rules or similarity calculations.
[0087] S60: Collect feedback information on the target project, and update the knowledge graph and the decision model based on the feedback information.
[0088] In this embodiment, the collection of feedback information on the target project is based on the observation of the user's actual interactive behavior, covering the user's click, browse, purchase, favorite, reject, and complain actions after the recommended project is displayed. This behavioral information needs to be structured and annotated in combination with the context state, including dimensions such as behavior type, behavior occurrence time, behavior duration, and terminal type. The processing of feedback information first involves behavioral attribution, that is, identifying whether the feedback is truly directly related to the current target project, and then screening out high-confidence data for model updates and knowledge graph corrections.
[0089] Feedback behaviors are categorized and mapped to feedback types. For example, positive feedback may include purchases and extended browsing, while negative feedback may include skipping, closing, and complaints. Structured instructions are generated based on the feedback type, identifying the relationships that need to be updated between the user entities involved and the target item entities, and defining update operation types, such as strengthening edge weights, adding behavior labels, and updating timestamps.
[0090] Knowledge graph updates are typically performed by adjusting entity edge weights, changing edge attributes, or adding new semantic paths. If users frequently provide positive feedback on products in a certain category, their connection weight in the graph can be dynamically increased to reflect their preferences.
[0091] In the decision-making model, feedback information is converted into training samples and fed back into the parameter update process of the policy network and value network through a reinforcement learning mechanism. The reward function maps different types of feedback into numerical reward values, which are used to calculate the reward deviation of the current policy for that state-action pair. In combination with policy optimization algorithms such as PPO and DDPG, parameter gradients are calculated and model weights are updated, gradually optimizing the model behavior.
[0092] On a health insurance push platform, a user receives a recommendation for a "Family Critical Illness Insurance Plan" and takes a positive action, clicking to view and completing the purchase. The system records the contextual features of this push (user profile, graph path, recommendation probability) and the interaction results, and labels the purchase action as high-weighted positive feedback. The platform then incorporates this feedback into the knowledge graph, increases the weight of the edge between the "User Entity → Family Insurance Product Entity" relationship, adds the "Purchase Action" label, and updates the last interaction time.
[0093] In parallel, the feedback system converts this sample into part of the training set and feeds it into the model optimization module. This data, by setting a high reward value, reinforces the rationality of the current policy's output of this action in this state. The policy network further strengthens the probability of selecting this state-action pair in the next round of training.
[0094] In another scenario, if a user closes three recommended financial product pages in a short period of time, the system interprets this as negative feedback, maps this behavior to a negative reward, and simultaneously reduces the weight of the connection to that product category in the knowledge graph. In the future, when similar situations occur, the recommendation weight of that product category will be lowered, effectively preventing users from feeling negatively impacted.
[0095] Example: In a healthcare service scenario, a user accepted the platform's recommended "Hypertension Chronic Disease Management Plan" and completed an online consultation and uploaded their initial blood pressure monitoring data within two days. This behavior was recorded as combined positive feedback, and the system mapped this behavior as a positive reinforcement path from "User → Chronic Disease Management Service" in the graph. The model also assigned a high reward value to this behavior. In the next round of policy training, the model was more inclined to recommend management and long-term intervention services to the user.
[0096] In a fintech business scenario, a user didn't click on the system's recommended "Smart Bond Portfolio" and instead clicked on "Non-Fixed Income Products" twice in a row. The system identified this combination of behaviors as a signal of a preference shift and adjusted the labels of the relevant feature dimensions in the user profile. The edge weight for "User → Fixed Income Product" in the knowledge graph was weakened, and a low reward was assigned to this state-action pair during model training. Ultimately, the model reduced this type of recommendation and, in the new state, outputted products that better matched the user's risk preferences.
[0097] This embodiment records and analyzes users' actual feedback on target items, converting this behavioral data into structured update instructions. This then strengthens or weakens the entity relationships in the knowledge graph and dynamically adjusts decision model parameters based on a reward mechanism, effectively enabling the system to adapt to user preferences. This mechanism establishes a closed-loop "recommendation-feedback-adjustment" process, enabling the model to self-optimize and dynamically adapt, improving the accuracy and timeliness of recommendations.
[0098] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a product decision optimization method, device, equipment and medium based on knowledge graph, including: obtaining and integrating user-related information, product-related information and environmental information to generate target information; constructing a knowledge graph based on the target information; constructing a user feature portrait based on the user-related information and the knowledge graph; using the user feature portrait and the information in the knowledge graph as input states, and the product information set as the action space, to train and generate a decision model based on a preset incentive mechanism; outputting the target project through the decision model; collecting feedback information on the target project, and updating the knowledge graph and decision model based on the feedback information. The present invention improves the comprehensiveness of the input state through a fusion modeling method that includes user information, product information and environmental information, combines the knowledge graph and user feedback to build a dynamic update mechanism, drives the continuous optimization of the decision model, and thereby improves the pertinence of the recommendation results and the adaptive ability of the system.
[0099] In one embodiment, the above step S10 includes:
[0100] S101, collecting multi-source heterogeneous data including user behavior log data, product attribute data, and environmental dynamic data;
[0101] S102, performing format standardization processing on the multi-source heterogeneous data, and unifying the time format and coding standard to generate standardized data;
[0102] S103, performing a data cleaning operation on the standardized data, detecting and processing missing values and outliers, and generating cleaned data;
[0103] S104, performing data structure conversion on the cleaned data to generate a structured data table;
[0104] S105: Integrate the structured data table to generate target information.
[0105] In this embodiment, the operation of collecting multi-source heterogeneous data usually involves different data sources from the user end, product database, and external environment monitoring system. There are differences in format, structure, and update frequency between these data sources. User behavior log data can be derived from records such as user click streams, page dwell time, and interaction paths on the platform. Product attribute data includes fields such as product name, classification label, price, applicable population, and historical evaluation. Environmental dynamic data includes but is not limited to external events such as holiday schedules, weather information, financial market fluctuations, and public health events. Data collection is usually achieved through multi-threaded pulling, interface docking, or scheduled task scheduling. It is necessary to ensure that the data integrity and collection frequency are reasonably matched.
[0106] Due to the heterogeneity of collected data, format standardization is necessary. This includes standardizing timestamps to a comparable format (such as ISO 8601) and unifying various coding methods (such as converting geographic location information to standard region codes and product IDs to platform-standard coding rules). This step can be achieved through ETL tools, regular expression parsing modules, or configurable data processing engines. Standardized data facilitates subsequent cleaning, conversion, and analysis, reducing data redundancy and processing errors.
[0107] On top of standardization, data cleaning operations are performed. These typically include filling in missing values (e.g., using strategies like the mean, median, or model prediction), detecting outliers (e.g., detecting atypical patterns in user behavior based on statistical distribution, rule-based, or machine learning methods), removing duplicate data (e.g., comparing MD5 summaries and ensuring primary key uniqueness), and visually verifying and tracking abnormal records through data quality monitoring mechanisms. The goal of cleaning operations is to improve the authenticity, consistency, and availability of data.
[0108] After cleaning, the data enters the structured conversion phase. Unlike the original log format or nested JSON format, structured conversion logically splits various fields into a two-dimensional table structure, forming a standard data table with user ID, product ID, and timestamp as the primary key. This process uses data mapping rules, field matching models, and semantic recognition algorithms to normalize the table structure for easy querying and modeling.
[0109] Ultimately, the generated structured data tables are unified and integrated into a single data representation object, the target information. This integration process requires identifying foreign key relationships, entity correspondences, and time alignment logic between different data tables. For example, user-product interactions can be linked to product IDs through user IDs, and external environmental events can be linked to user behavior records through time. The target information should be complete, traceable, and cover the feature dimensions required for modeling. It serves as the core input foundation for subsequent operations such as knowledge graph construction and model training.
[0110] This embodiment collects multi-source heterogeneous data from three dimensions: user behavior, product attributes, and environmental context, and performs standardization, cleansing, and structuring on the data. This effectively breaks down information silos and builds a unified data view. The resulting integrated target information not only has unified data semantics and formatting, but also exhibits good temporal consistency and feature dimension completeness. This provides higher data quality support for subsequent knowledge graph construction and recommendation model training, significantly improving the system's adaptability and usability.
[0111] In one embodiment, the above step S20 includes:
[0112] S201, identifying a user entity and a product entity from the target information, and extracting entity attribute information from the target information to generate an entity set including entities and their attributes;
[0113] S202, analyzing the interaction pattern between entities in the target information, determining the association relationship between the user entity and the product entity, and generating a relationship set;
[0114] S203, using entities in the entity set as nodes and relationships in the relationship set as edges to generate a knowledge graph, and storing it in a graph database;
[0115] S204: When new target information is added, the entity nodes and relationship edges in the knowledge graph are dynamically updated.
[0116] In this embodiment, the process of identifying user entities and product entities from the target information first requires performing entity extraction operations based on user identification fields (such as user ID, account name, mobile phone number) and product identification fields (such as product ID, SKU code, product name). This process is generally completed through an entity recognition module, which can perform high-precision extraction based on field mapping rules, regular template matching, data dictionary support or pre-trained models (such as NER models). The extracted user entities and product entities must be unique, and homologous objects can be merged through primary key deduplication or entity fusion methods, and standardized named entity labels can be generated.
[0117] Extracting entity attribute information from the target information typically involves parsing the metadata corresponding to each field in a structured table. For example, for a user entity, attributes such as age, gender, region, income level, and historical purchase frequency can be extracted; for a product entity, attributes such as product type, price range, coverage, and target demographics can be extracted. These attributes are attached to the corresponding entity as key-value pairs, forming a structure that maps the primary entity to the attribute fields, forming a complete entity set.
[0118] Analyzing the interaction patterns between entities in the target information primarily relies on user-product behavior records, such as clicks, browsing, purchases, reviews, and policy cancellations. By identifying contextual information such as the time, frequency, and sequence of these behaviors, semantic connections between users and products can be constructed, such as logical edges like "purchased - May 2024" and "multiple views - no purchase." These edges form a set of relationships, each of which includes not only the start and end entity identifiers of the connection but also metadata such as the type of interaction, timestamp, and intensity, enhancing the graph's expressive power.
[0119] Mapping entity and relationship sets into a graph database structure requires a database engine that supports property graphs, such as Neo4j, JanusGraph, or TinkerGraph. In this graph structure, user and product entities are represented as nodes, each containing its attribute fields. Interaction relationships are represented as edges connecting the nodes, carrying interaction labels and contextual attributes. This graph structure should include indexing mechanisms, query interfaces, and concurrency support to ensure efficient access and reasoning.
[0120] When adding new target information, the incremental update mechanism needs to be triggered. By detecting changes in user identifiers, product identifiers, and behavior records in the new data, it is determined whether it is a new entity or a new relationship. If it is a new entity, a new node is added; if it is an existing entity but the attribute has changed, the attribute value is updated; if it is a new interaction record, a new edge is added. The incremental update process requires maintenance. Figure 1 Consistency is maintained to avoid entity duplication, edge conflicts, or attribute loss, while recording update logs for auditing and backtracking.
[0121] This embodiment achieves a structured and highly semantic knowledge representation by identifying entities, extracting attributes, parsing interaction patterns, and mapping them to a graph structure from target information. This knowledge graph not only connects the logical relationship between users and products, but also embeds rich attribute information and dynamic interaction context, providing strong knowledge support for subsequent recommendation strategies. The continuous incremental update capability of the graph database ensures that the knowledge graph responds quickly to data changes, giving the system high scalability and semantic reasoning capabilities, thereby improving the contextual relevance and personalized adaptation of recommendation results.
[0122] In one embodiment, the above step S30 includes:
[0123] S301, parsing user basic attribute data from the user-related information to generate a basic attribute feature set;
[0124] S302, querying the product interaction records associated with the user entity in the knowledge graph to generate a knowledge graph feature set;
[0125] S303, fusing the basic attribute feature set and the knowledge graph feature set to generate a fused feature vector;
[0126] S304: Perform feature encoding processing on the fused feature vector to generate a user feature profile.
[0127] In this embodiment, parsing user basic attribute data from user-related information typically relies on raw user data provided by the data access layer, including registration information, real-name information, authentication information, or feature data derived from historical transaction behavior. Core fields in this information, such as age, gender, region, marital status, occupation type, annual income level, and social security status, can be considered static attributes with high stability, low timeliness, but strong discrimination. In practical applications, the required fields can be extracted from the user database through field mapping rules and combined with a standardized coding system to form a structured basic attribute feature set, which serves as the static basis for subsequent feature expression.
[0128] Querying product interaction records associated with user entities in the knowledge graph relies on querying the adjacency of user nodes in the graph database. This typically involves reading all interaction edges connected to the user node and the target product node. These edge types can include browsing, clicking, purchasing, commenting, surrendering, and claiming, and each type may also be accompanied by contextual information such as timestamps, behavior frequency, and result status. To enhance expressiveness, a scoring mechanism based on behavior weights can also be introduced to assign different influences to different interaction behaviors. For example, "purchase" may have a higher weight than "browse," thereby generating a knowledge graph feature set that better reflects users' dynamic preferences.
[0129] To fuse the basic attribute feature set with the knowledge graph feature set, vectorized encoding methods are used to unify the two feature types into the same representation space. Basic attribute features can be represented using one-hot encoding, binning encoding, or vector embedding; graph features can be represented using path encoding, edge type embedding, aggregated neighbor node representation, and other methods. Fusion strategies such as vector concatenation, weighted summation, or attention mechanisms can be used to generate the fused feature vector, ensuring compatible representation of static and dynamic information and trainability.
[0130] Feature encoding is performed on the fused feature vector, requiring vector normalization, dimensionality scaling, and noise reduction to ensure that the vector meets model input specifications. This process can be implemented using standard machine learning preprocessing modules (such as the Scaler series in Sklearn) or embedding layers in neural networks. Positional encoding, multi-head attention mechanisms, or contextual encoding modules can also be introduced to enhance semantic relevance. The output vector after encoding is a user feature profile, expressive enough to be directly used by downstream models.
[0131] This embodiment integrates the user's static attributes with their dynamic behavior in the knowledge graph to construct a user representation vector that is both stable and timely, fully reflecting the user's current interests, preferences, and historical behavior patterns. This representation method has strong semantic recognition and generalization capabilities, and can adapt to the input requirements of multiple recommendation models. At the same time, through feature encoding processing, the numerical stability and semantic consistency of the vector are improved, providing a high-quality, high-dimensional, low-redundancy input foundation for subsequent modeling, thereby significantly enhancing the accuracy and flexibility of the system's personalized recommendations.
[0132] In one embodiment, the above step S40 includes:
[0133] S401, combining the user feature profile and the information in the knowledge graph into a state vector;
[0134] S402, defining a product information set as a discrete action space;
[0135] S403, constructing a policy network with the state vector as input and the probability distribution of each action in the discrete action space as output;
[0136] S404, constructing a value network with the state vector and action as input and the value of the state-action combination as output;
[0137] S405, designing a reward function that allocates reward values according to user feedback types;
[0138] S406 , using a policy optimization module, the parameters of the policy network and the value network are updated in combination with the reward function to generate a decision model.
[0139] In this embodiment, combining the user profile with information from the knowledge graph into a state vector means integrating static features and dynamic context into a unified data structure, forming an environmental state representation for reinforcement learning modeling. The user profile portion consists of vectorized user static attributes and behavioral features, while the knowledge graph information portion includes the semantic feature embedding results of user-associated nodes, structural information about the relationships between nodes, and edge weight distribution. These two parts can be constructed into a unified state vector through splicing or fusion layers, maintaining the model's comprehensive perception of the user's overall preferences and scene context.
[0140] Define the product information set as a discrete action space. All recommended products must be discretely numbered or indexed, with each number corresponding to a specific product or set of products, forming the action dimension in reinforcement learning. This action space must be enumerable and identifiable, and consistent with the product resource management structure in actual recommendation systems, ensuring accurate mapping to recommended entities in real-world deployments.
[0141] A policy network is constructed that takes a state vector as input and outputs the probability distribution of each action in a discrete action space. This network is primarily used to generate a recommendation intent distribution. This policy network can be constructed using a multilayer perceptron (MLP), graph neural network (GNN), or attention network. Its output is the selection probability of all possible products, representing the model's behavioral tendency in the current state.
[0142] A value network, which takes a state vector and an action as input and outputs the value of the state-action combination, is a crucial component in reinforcement learning for evaluating the effectiveness of current choices. The value network accepts a combination of state and action as input and outputs a scalar value that measures the expected long-term benefit of executing that action under the current state. Common implementations include a two-stream neural network architecture or a distributed valuation method based on temporal difference learning.
[0143] Design a reward function that assigns rewards based on the type of user feedback. This requires mapping different types of user behavior to specific feedback signals. For example, completing a purchase could be given a highly weighted positive reward, browsing could be given a less weighted positive reward, and rejecting a request could be given a negative reward. The reward function can be flexibly adjusted using strategies such as linear mapping, weight decay, or reward differentials, supporting strategy tuning during dynamic learning.
[0144] The policy optimization module combines the reward function to update the policy network and value network parameters, and iterative training is performed using reinforcement learning methods such as policy gradient methods, deep Q-learning (DQN), and proximal policy optimization (PPO). The policy optimization module should be able to continuously receive environmental feedback, process gradient information, and update model parameters, allowing the network to converge to the optimal policy space. Ultimately, it trains a decision model with high dynamic responsiveness and the ability to adapt to user preferences.
[0145] This embodiment realizes an adaptive learning mechanism for a personalized recommendation system by fusing user feature profiles with knowledge graph information to form a state expression and constructing a training structure based on the collaborative optimization of reinforcement learning policy networks and value networks. After introducing a reward function driven by behavioral feedback, the model can not only summarize behavioral preferences from user history, but also dynamically adjust policy parameters based on real-time feedback, thereby continuously optimizing the accuracy of recommendation decisions in a complex and changing recommendation environment. This mechanism effectively improves the system's generalization ability and response efficiency when facing diverse user and product scenarios, making recommendation results more targeted and timely.
[0146] In one embodiment, the above step S50 includes:
[0147] S501, obtaining the current user feature profile and the current knowledge graph information, and combining them into an inference state vector;
[0148] S502, inputting the inference state vector into the decision model to generate a selection probability distribution of each product in the action space;
[0149] S503, performing probability sampling according to the selection probability distribution to determine the target project identifier;
[0150] S504: Convert the target project identifier into a target project.
[0151] In this embodiment, obtaining the current user profile and current knowledge graph information and combining them into an inference state vector is an integration operation on the current context information before performing recommendation inference. The user profile is derived from the encoded representation of the user's static attributes (such as age, region, and occupation) and dynamic behavioral characteristics (such as historical clicks and purchase frequency). The knowledge graph information is derived from the product nodes, relationship edges, and attribute values directly or indirectly connected to the user entity in the graph structure. These are extracted through graph neural network embedding or structural encoding methods. The two are spliced together to form a representation of the current environment state, ensuring that the association between individual preferences and knowledge structure is considered during the inference process.
[0152] Inputting this inference state vector into the decision model executes a forward propagation operation. The model, having previously trained its parameters to optimize, is capable of generating a distribution of recommendation preferences for different state inputs. The model typically includes a policy network, whose output is a vector corresponding to the number of recommended products, with each dimension representing the probability of selecting that product under the current state. This output vector can be understood as a probability distribution for each action in the action space, providing a probabilistic basis for subsequent decisions on the target item.
[0153] Probabilistic sampling based on a selection probability distribution is a method for selecting specific items from the policy output when generating recommendations. This operation involves performing a random sampling in the action space, where the probability of each item being sampled is equal to its corresponding value in the policy distribution. This method avoids the overfitting problem of single-maximum selection, allowing for exploratory and diverse recommendation results. It is particularly suitable for scenarios where recommended content is frequently updated or user preferences are unclear.
[0154] Converting target item identifiers to target items is the process of mapping the discrete identifiers output by the model to actual product entities for recommendation. This identifier is typically a unique index number within a product information collection. Through a table lookup or mapping mechanism, it is converted into a specific product information entry, including attributes such as name, description, price, warranty coverage, or applicable scenarios, for use in the final recommendation display and push service.
[0155] This embodiment fuses user feature profiles with knowledge graph information to form an inference state vector, and combines it with a decision model for probabilistic reasoning and sampling operations, enabling the recommendation system to output personalized products based on individual differences. This mechanism not only ensures a high degree of match between recommendation results and user interests, but also enhances the diversity and exploratory nature of recommended content by introducing probabilistic sampling methods. In a dynamic environment, this method can continuously adapt to changes in user needs, improve the system's recommendation accuracy and response intelligence, and effectively support precise matching and recommendation output in intelligent service scenarios in areas such as medical insurance and financial products.
[0156] In one embodiment, the above step S60 includes:
[0157] S601, collecting user feedback behavior data on the target project;
[0158] S602, analyzing the feedback behavior data to determine whether the feedback type is a purchase behavior, a browsing behavior, or a rejection behavior;
[0159] S603, generating a knowledge graph update instruction based on the feedback type;
[0160] S604, updating the entity nodes and relationship edges in the knowledge graph according to the knowledge graph update instruction;
[0161] S605, determining a reward value based on the feedback type and a preset incentive mechanism, and determining a model update parameter based on the reward value;
[0162] S606: Update the network weight of the decision model based on the model update parameter.
[0163] In this embodiment, collecting user feedback behavior data on the target project refers to real-time or periodic collection of user behavior performance in the system after the target project is pushed. Common data forms include click records, dwell time, purchase conversion, page bounces, collection behavior or explicit rejection operations. These behaviors can be obtained through client-side tracking, server-side logs, API feedback, etc., to form first-hand data on the acceptance of recommendation results.
[0164] Analyzing feedback behavior data to determine the type of feedback typically relies on rule-based mapping or classification models, primarily categorizing raw behavior data into three types: "purchase behavior," "browsing behavior," or "rejection behavior." For example, completing an order can be mapped as a purchase behavior, merely opening a page without further interaction can be mapped as a browse behavior, and explicitly clicking "not interested" or quickly exiting the page can be classified as a rejection behavior. This classification directly influences subsequent adjustments to the graph and model.
[0165] Knowledge graph update instructions are generated based on the feedback type, aiming to dynamically adjust the graph database structure based on the actual user interaction results. If a purchase behavior is identified, an instruction to add a "Purchase" relationship edge may be generated; if a rejection behavior is identified, a negative preference flag or relationship weight decay instruction may be generated. These instructions describe the operation content and attribute fields that should be added, deleted, or modified between the user and product entities in the graph, ensuring that the graph always reflects the latest user preference trends.
[0166] Updating the entity nodes and relationship edges in the knowledge graph according to the above instructions is a data structure operation performed on the graph database. For example, with graph databases such as Neo4j, you can use Cypher statements to perform operations such as CREATE, MERGE, or SET to add new edges, adjust attribute weights, or mark preferred directions based on the original graph structure, thereby achieving structured incremental updates of the graph.
[0167] Determining reward values based on feedback type and a pre-set incentive mechanism quantifies user feedback signals into numerical metrics to guide model learning. Incentive mechanisms typically predefine reward values for different types of feedback, such as assigning positive high rewards to purchases, neutral low rewards to browsing, and negative penalties to rejections. Reward values reflect the positive or negative contribution of a recommendation result to the overall strategy and constitute the core source of training signals.
[0168] Based on this reward value, model update parameters are further determined, typically including the Temporal Difference Error (TD error) in the policy gradient algorithm, the adjustment range of the target network weights, or the learning rate adjustment coefficient. This process combines the difference between the current policy output and the feedback reward to calculate the loss function and perform backpropagation, extracting gradient information for optimizing model weights.
[0169] Finally, updating the decision model's network weights based on the model update parameters involves performing one or more model parameter adjustments. Using a neural network optimizer (such as Adam or RMSprop), the policy and value network parameters are gradient-updated, increasing their preference for actions with high reward probabilities in future recommendations. This model update mechanism provides continuous self-adaptation, enabling continuous adjustments to recommendation preferences in dynamic scenarios and enhancing the system's online learning capabilities.
[0170] For example: widely collect customer data through multiple channels, including basic information (age, gender, occupation, income, etc.), behavioral data (APP browsing records, consultation records, purchase history, etc.), and external market data (industry dynamics, policies and regulations, economic indicators, etc.). Use data cleaning and conversion technology to integrate data in different formats and sources to prepare for subsequent processing. Based on the integrated data, determine entities such as customers, insurance products, insurance terms, industry terms, and the relationships between them (such as the purchase relationship between customers and products, the inclusion relationship between products and terms, etc.). Use entity extraction, relationship extraction and other technologies to extract knowledge from structured, semi-structured and unstructured data, and build a dynamic knowledge graph through entity alignment and knowledge fusion. The knowledge graph can be updated in real time to reflect the latest customer information, product information and market dynamics.
[0171] Utilizing message queue technologies (such as Kafka), data from multiple sources is collected in real time, including customer app operation logs, transaction data generated by business systems, and industry trends pushed by external information platforms. These data sources continuously enter the system, providing material for updating the knowledge graph. Next, stream computing frameworks (such as Apache Flink) are used to clean the collected raw data in real time, removing noise, filling missing values, and unifying the data format. For example, time data in different formats can be unified into a standard time format to ensure data accuracy and availability, laying the foundation for subsequent knowledge extraction. For structured data, SQL queries are written to directly extract entity information such as customers and insurance products, along with their attribute information, from database tables. For example, information such as a customer's age, gender, and contact information can be obtained from a customer information table. For semi-structured data, such as insurance terms documents in XML or JSON format, specific parsers are used to extract key information and determine the relationships between terms and products. For unstructured data such as customer consultation records and news reports, deep learning-based named entity recognition (NER) models, such as those based on the BERT architecture, are used to identify entities within the text, such as the names of diseases and insurance products mentioned by customers. Relationship extraction techniques, such as those based on graph convolutional networks (GCNs), are then used to analyze the semantic relationships between entities within the text and determine connections between customers and products, products and terms, and so on. The goal of the relationship extraction model is to maximize the conditional probability P(R|E_1,E_2,T), where R is the relationship, E_1 and E_2 are two entities, and T is the text. The model is trained on a large amount of annotated data to learn the mapping between text features and relationships. Neo4j, a graph database, is used to store the knowledge graph. Neo4j represents entities using nodes and relationships using edges, enabling efficient storage and querying of complex connected data. Once new entities or relationships are extracted, insertion or update operations are performed using Neo4j's Cypher query language. For example, when a new customer is found to have purchased a certain insurance product, a new "purchase" relationship edge is generated to connect the customer node and the product node, and relevant attributes such as purchase time, premium amount, etc. are updated. At the same time, a timestamp mechanism is established to record the time of each knowledge graph update for easy traceability and management. Utilizing the graph structure and existing knowledge of the knowledge graph, potential knowledge is mined through rule reasoning and machine learning-based reasoning methods. Rule reasoning is based on pre-set business rules, such as "If a customer purchases car insurance, then the customer may need to purchase accident insurance", and is matched and deduced in the knowledge graph. Machine learning-based reasoning uses the features and relationships of nodes to train graph neural network models (such as GraphSAGE) to predict the potential relationship between customers and products, providing more comprehensive knowledge support for recommendation decisions. The GraphSAGE model generates node representations by aggregating the features of neighboring nodes. The representation h of node vv Calculated by the following formula:
[0172]
[0173] The representation vector (embedding) representing node v at the kth layer is the output after the aggregation operation of the kth layer graph neural network. Represents the representation vector of node v in the k-1th layer, that is, the feature representation of the current layer input itself. Represents the representation vector of node u (i.e., neighbor node) at the k-1th layer. N(v) represents the set of neighbor nodes of node v. It is used to aggregate the features of neighbor nodes. Indicates the average pooling (meanaggregation) of the features of neighboring nodes, that is, the summary expression of neighbor information. CONCAT(·,·) represents the connection operation, which combines the node's own representation The aggregated representations of the W and its neighbors are concatenated together as the combined input. k The weight matrix of the kth layer is a parameter that needs to be learned through training. It is used to perform a linear transformation on the concatenated feature vector. σ represents a nonlinear activation function, such as ReLU, sigmoid, and tanh, which is used to introduce nonlinear feature expression capabilities.
[0174] Based on customer data and information from the knowledge graph, we use techniques such as feature extraction and cluster analysis to construct customer profiles. These profiles cover multiple dimensions, including basic customer attributes, behavioral characteristics, insurance needs, and risk preferences. As customer behavior and information change, these profiles are updated in real time to ensure accuracy and timeliness.
[0175] The customer profile and dynamic knowledge graph information are used as the input state of the reinforcement learning model, and insurance products are used as the action space. A reasonable reward function is designed to provide corresponding rewards based on customer feedback on recommended products (purchase, browsing, consultation, etc.). Through continuous interaction with the environment, reinforcement learning algorithms such as the proximal policy optimization algorithm are used to train the policy network and value network, optimizing the recommendation strategy to maximize long-term cumulative rewards. The details are as follows:
[0176] The reinforcement learning model integrates multi-dimensional information from the customer profile, such as age, income, list of purchased insurance products, and recent product browsing history, along with customer-related information extracted from the dynamic knowledge graph, such as the customer's risk group and potential product associations. The various products in the insurance product library, such as critical illness insurance, medical insurance, and accident insurance, constitute the action space.
[0177] Use a neural network based on the Transformer architecture to build a policy network π θ(a t ∣s t ), indicating that in state s t Next select action a t The probability distribution of is controlled by the network parameters θ. θ refers to all trainable parameters in the policy network, including attention weights, feedforward network parameters, etc. The multi-head attention mechanism in the Transformer architecture can effectively capture the relationship between different features in the input state, and the fully connected layer further fuses and transforms the output of the attention mechanism. The policy network takes the state vector as input and outputs the recommended probability distribution of each action (insurance product) through a series of linear transformations and activation function operations. In the multi-head attention mechanism, for the input feature x, the formula for calculating the attention score is:
[0178]
[0179] Where Q, K, and V are query, key, and value matrices, respectively, which are obtained by linearly transforming the input feature x, that is, Q = xW Q , K=xW K , V=xW V , W Q 、W K 、W V is the learnable weight matrix, d k is the dimension of the key vector. Multi-head attention concatenates the results of multiple attention heads and linearly transforms them again, namely:
[0180] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0181] in h is the number of attention heads, W O These are all learnable weight matrices. The model selects specific insurance products for recommendation based on probability distribution and an exploration-exploitation strategy.
[0182] Also based on the Transformer architecture to build the value network V ω (s t ,a t ), where w is the network parameter. Its input is the state s t and action a t The value network estimates the current state s by learning t Take an action a tAfter a series of calculations in the Transformer network (including multi-head attention calculations and fully connected layer processing, the calculation process is similar to the policy network), a scalar value V is finally output. ω (s t ,a t ), which represents the value of the state-action pair. The value network provides an evaluation basis for the decision-making of the policy network, helping the policy network optimize the recommendation strategy.
[0183] When a customer purchases a recommended insurance product, a larger positive reward is given. Reward value R buy It can be determined based on factors such as the product's premium amount p, profit margin pr, and customer long-term value v. The formula is R buy =αp+βpr+γv, where α, β, and γ are weight coefficients determined through experiments or business experience and used to balance the impact of various factors on rewards.
[0184] If customers browse the recommended products, they will be given a certain positive reward R view , the reward value is lower than the purchase reward and is usually set to a fixed value, such as R view =δ, which aims to encourage the model to explore products of potential interest to customers.
[0185] When customers consult for recommended products, they are given positive rewards R inquire , also set to a fixed value, such as R inquire =μ, indicating that the customer has a certain interest in the product.
[0186] If the recommended product does not match the customer profile seriously, or the customer shows aversion to the recommendation (such as complaining, quickly closing the recommendation page), a negative reward R is given. negative = -v, where v is a positive number that encourages the model to avoid such bad recommendations.
[0187] Taking all situations into consideration, the reward function R(s t ,a t ) can be expressed as R(s t ,a t )=R buy +R view +R inquire +R negative (Only the corresponding reward items will be activated according to the actual situation.) For example, if a customer browses the recommended product but does not buy it or consult, then R(s t ,a t )=R view .
[0188] The proximal policy optimization algorithm (PPO) is used to train the policy network and value network. During the training process, first, the current policy π θ Interact with the environment (customer) and collect a series of states s t 、Action a t , reward r t and the next state s t+1 The advantage function A(s) is calculated by generalized advantage estimation (GAE) t ,a t ), which is used to measure the advantage of the current action relative to the average action value. The calculation formula of GAE is:
[0189]
[0190] Where γ is the discount factor used to measure the importance of future rewards, and its value range is 0≤γ≤1. λ is the GAE parameter, and its value range is 0≤λ≤1. t+k =r t+k +γV ω (s t+k+1 )-V ω (s t+k ) is the TD error. According to the objective function of the PPO algorithm, the policy network parameters θ are updated by maximizing the product of the ratio of the action probability output by the policy network and the action probability under the old policy and the advantage function. Objective function J PPO (θ) is:
[0191]
[0192] where π θold are the parameters of the old policy network. To prevent the policy update from being too large, a pruning mechanism is introduced. The actual optimization goal is:
[0193]
[0194] Where clip is the clipping function and ∈ is the clipping parameter. By limiting the amplitude of the policy update, the stability of the policy is guaranteed. At the same time, the value network parameter ω is updated according to the loss function of the value network, such as the mean square error loss function L(ω). The loss function is:
[0195]
[0196] Using optimization algorithms like stochastic gradient descent, we calculate the gradient of the loss function and update the value network parameters to more accurately estimate the value of state-action pairs. Through continuous iterative training, we improve the recommendation performance of the reinforcement learning model.
[0197] Based on a trained reinforcement learning model, the most suitable insurance products are recommended to customers in real time. Customer feedback on recommended products (such as purchase, rejection, further consultation, etc.) is collected and used to update the dynamic knowledge graph and reinforcement learning model, forming a closed-loop optimization process to continuously improve the accuracy and effectiveness of recommendations.
[0198] This embodiment collects real user feedback on recommended items and divides the feedback behavior into multiple types, driving the two-way dynamic update of the knowledge graph and the policy model, so that the system can continuously optimize the recommendation path based on the actual user behavior. On the one hand, the structural changes in the knowledge graph reflect the latest migration trends of user preferences and enhance the expressiveness between entities; on the other hand, the online optimization of the decision model parameters ensures the matching efficiency between the policy output and the target response. This mechanism realizes the closed-loop self-evolution of the recommendation system, improves the relevance of the recommended content and the accuracy of the response. It is especially suitable for scenarios with high requirements for personality differences, behavioral changes and multi-target adaptability, such as medical insurance recommendations and precise delivery of financial products.
[0199] In one embodiment, a product decision optimization device based on a knowledge graph is provided, and the product decision optimization device based on a knowledge graph corresponds one-to-one to the product decision optimization method based on a knowledge graph in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of the knowledge graph-based product decision optimization device of the present invention. It includes a target information generation module 10, a graph construction module 20, a user profile construction module 30, a decision model training module 40, a target item inference module 50, and a model update module 60. Each functional module is described in detail below:
[0200] The target information generation module 10 is used to obtain and integrate user-related information, product-related information, and environmental information to generate target information;
[0201] A graph construction module 20, configured to construct a knowledge graph based on the target information;
[0202] A user portrait construction module 30 is used to construct a user feature portrait based on the user-related information and the knowledge graph;
[0203] A decision model training module 40 is configured to train and generate a decision model based on a preset incentive mechanism, using the user feature profile and the information in the knowledge graph as input states and the product information set as an action space;
[0204] A target item reasoning module 50 is configured to output a target item through the decision model;
[0205] The model updating module 60 is used to collect feedback information on the target project and update the knowledge graph and the decision model based on the feedback information.
[0206] In one embodiment, the target information generating module 10 is specifically configured to:
[0207] Collect multi-source heterogeneous data including user behavior log data, product attribute data, and environmental dynamic data;
[0208] Performing format standardization processing on the multi-source heterogeneous data, and unifying the time format and coding standard to generate standardized data;
[0209] Performing data cleaning operations on the standardized data, detecting and processing missing values and outliers, and generating cleaned data;
[0210] Performing data structuring conversion on the cleaned data to generate a structured data table;
[0211] The structured data tables are integrated to generate target information.
[0212] In one embodiment, the graph construction module 20 is specifically configured to:
[0213] Identifying user entities and product entities from the target information, extracting entity attribute information from the target information, and generating an entity set including entities and their attributes;
[0214] Analyze the interaction patterns between entities in the target information, determine the association relationship between user entities and product entities, and generate a relationship set;
[0215] Using the entities in the entity set as nodes and the relationships in the relationship set as edges, a knowledge graph is generated and stored in a graph database;
[0216] When new target information is added, the entity nodes and relationship edges in the knowledge graph are dynamically updated.
[0217] In one embodiment, the user portrait construction module 30 is specifically configured to:
[0218] Parsing user basic attribute data from the user-related information to generate a basic attribute feature set;
[0219] Querying the product interaction records associated with the user entity in the knowledge graph to generate a knowledge graph feature set;
[0220] Fusing the basic attribute feature set and the knowledge graph feature set to generate a fused feature vector;
[0221] Perform feature encoding processing on the fused feature vector to generate a user feature portrait.
[0222] In one embodiment, the decision model training module 40 is specifically configured to:
[0223] Combining the user feature profile and the information in the knowledge graph into a state vector;
[0224] Define the product information set as a discrete action space;
[0225] Constructing a policy network that takes the state vector as input and outputs the probability distribution of each action in the discrete action space;
[0226] Constructing a value network with the state vector and action as input and the value of the state-action combination as output;
[0227] Design a reward function that assigns reward values based on the type of user feedback;
[0228] The policy optimization module is used to update the parameters of the policy network and the value network in combination with the reward function to generate a decision model.
[0229] In one embodiment, the target item reasoning module 50 is specifically configured to:
[0230] Obtain the current user feature profile and current knowledge graph information and combine them into an inference state vector;
[0231] Inputting the inference state vector into the decision model to generate a selection probability distribution of each product in the action space;
[0232] Performing probability sampling according to the selection probability distribution to determine the target project identifier;
[0233] The target project identifier is converted into a target project.
[0234] In one embodiment, the model updating module 60 is specifically configured to:
[0235] Collect user feedback behavior data on the target project;
[0236] Analyze the feedback behavior data to determine whether the feedback type is a purchase behavior, a browsing behavior, or a rejection behavior;
[0237] Generate a knowledge graph update instruction based on the feedback type;
[0238] Update the entity nodes and relationship edges in the knowledge graph according to the knowledge graph update instruction;
[0239] Determining a reward value based on the feedback type and a preset incentive mechanism, and determining a model update parameter based on the reward value;
[0240] The network weights of the decision model are updated based on the model update parameters.
[0241] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the service side of a product decision optimization method based on a knowledge graph.
[0242] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. Among them, the processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a product decision optimization method based on a knowledge graph.
[0243] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0244] Obtain and integrate user-related information, product-related information, and environmental information to generate target information;
[0245] Constructing a knowledge graph based on the target information;
[0246] Constructing a user feature profile based on the user-related information and the knowledge graph;
[0247] The user feature profile and the information in the knowledge graph are used as input states, the product information set is used as the action space, and a decision model is trained and generated based on a preset incentive mechanism;
[0248] outputting a target project through the decision model;
[0249] Collect feedback information on the target project, and update the knowledge graph and the decision model based on the feedback information.
[0250] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0251] Obtain and integrate user-related information, product-related information, and environmental information to generate target information;
[0252] Constructing a knowledge graph based on the target information;
[0253] Constructing a user feature profile based on the user-related information and the knowledge graph;
[0254] The user feature profile and the information in the knowledge graph are used as input states, the product information set is used as the action space, and a decision model is trained and generated based on a preset incentive mechanism;
[0255] outputting a target project through the decision model;
[0256] Collect feedback information on the target project, and update the knowledge graph and the decision model based on the feedback information.
[0257] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0258] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0259] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0260] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A product decision optimization method based on knowledge graph, characterized in that: The following steps are involved: Obtain and integrate user-related information, product-related information, and environmental information to generate target information; Constructing a knowledge graph based on the target information; Constructing a user feature profile based on the user-related information and the knowledge graph; The user feature profile and the information in the knowledge graph are used as input states, the product information set is used as the action space, and a decision model is trained and generated based on a preset incentive mechanism; outputting a target project through the decision model; Collect feedback information on the target project, and update the knowledge graph and the decision model based on the feedback information.
2. The product decision optimization method based on knowledge graph according to claim 1, characterized in that: Obtain and integrate user-related information, product-related information, and environmental information to generate target information, including: Collect multi-source heterogeneous data including user behavior log data, product attribute data, and environmental dynamic data; Performing format standardization processing on the multi-source heterogeneous data, and unifying the time format and coding standard to generate standardized data; Performing data cleaning operations on the standardized data, detecting and processing missing values and outliers, and generating cleaned data; Performing data structuring conversion on the cleaned data to generate a structured data table; The structured data tables are integrated to generate target information.
3. The product decision optimization method based on knowledge graph according to claim 1, characterized in that: Constructing a knowledge graph based on the target information, including: Identifying user entities and product entities from the target information, extracting entity attribute information from the target information, and generating an entity set including entities and their attributes; Analyze the interaction patterns between entities in the target information, determine the association relationship between user entities and product entities, and generate a relationship set; Using the entities in the entity set as nodes and the relationships in the relationship set as edges, a knowledge graph is generated and stored in a graph database; When new target information is added, the entity nodes and relationship edges in the knowledge graph are dynamically updated.
4. The product decision optimization method based on knowledge graph according to claim 1, characterized in that: Constructing a user feature profile based on the user-related information and the knowledge graph, including: Parsing user basic attribute data from the user-related information to generate a basic attribute feature set; Querying the product interaction records associated with the user entity in the knowledge graph to generate a knowledge graph feature set; Fusing the basic attribute feature set and the knowledge graph feature set to generate a fused feature vector; Perform feature encoding processing on the fused feature vector to generate a user feature portrait.
5. The product decision optimization method based on knowledge graph according to claim 1, characterized in that: The user feature profile and the information in the knowledge graph are used as input states, the product information set is used as the action space, and a decision model is trained and generated based on a preset incentive mechanism, including: Combining the user feature profile and the information in the knowledge graph into a state vector; Define the product information set as a discrete action space; Constructing a policy network that takes the state vector as input and outputs the probability distribution of each action in the discrete action space; Constructing a value network with the state vector and action as input and the value of the state-action combination as output; Design a reward function that assigns reward values based on the type of user feedback; The policy optimization module is used to update the parameters of the policy network and the value network in combination with the reward function to generate a decision model.
6. The product decision optimization method based on knowledge graph according to claim 1, characterized in that: Outputting target items through the decision model includes: Obtain the current user feature profile and current knowledge graph information and combine them into an inference state vector; Inputting the inference state vector into the decision model to generate a selection probability distribution of each product in the action space; Performing probability sampling according to the selection probability distribution to determine the target project identifier; The target project identifier is converted into a target project.
7. The product decision optimization method based on knowledge graph according to claim 1, characterized in that: Collecting feedback information on the target project and updating the knowledge graph and the decision model based on the feedback information, including: Collect user feedback behavior data on the target project; Analyze the feedback behavior data to determine whether the feedback type is a purchase behavior, a browsing behavior, or a rejection behavior; Generate a knowledge graph update instruction based on the feedback type; Update the entity nodes and relationship edges in the knowledge graph according to the knowledge graph update instruction; Determining a reward value based on the feedback type and a preset incentive mechanism, and determining a model update parameter based on the reward value; The network weights of the decision model are updated based on the model update parameters.
8. A product decision optimization device based on knowledge graph, characterized in that: The product decision optimization device based on knowledge graph includes: Target information generation module, used to obtain and integrate user-related information, product-related information and environmental information to generate target information; A graph construction module, used to construct a knowledge graph based on the target information; A user portrait construction module is used to construct a user feature portrait based on the user-related information and the knowledge graph; A decision model training module is used to train and generate a decision model based on a preset incentive mechanism, taking the user feature profile and the information in the knowledge graph as input states and the product information set as an action space; A target item reasoning module, configured to output a target item through the decision model; A model updating module is used to collect feedback information on the target project and update the knowledge graph and the decision model based on the feedback information.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a knowledge graph-based product decision optimization program stored in the memory and capable of running on the processor. When the knowledge graph-based product decision optimization program is executed by the processor, the steps of the knowledge graph-based product decision optimization method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores a product decision optimization program based on a knowledge graph, and when the product decision optimization program based on a knowledge graph is executed by a processor, the steps of the product decision optimization method based on a knowledge graph as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-modal knowledge-driven production equipment system multi-target operation mode optimization method
CN120875190A
Multimodal knowledge-driven production plant system multi-objective operating mode optimization method
CN120875190B
Business handling reservation state query method based on knowledge graph
CN121412395A
A business handling reservation state query method based on a knowledge graph
CN121412395B
Multi-agent collaborative product closed-loop optimization system and method based on user feedback data
CN121615880A