Security information personalized recommendation method and system based on knowledge graph
By constructing a multi-source knowledge graph and causal inference paths, the problems of dynamic changes in user interests and the impact of market events in securities information recommendation are solved, achieving more accurate and personalized securities information recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN JINHUI RONGZHI DATA SERVICE CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-07
Smart Images

Figure CN122347459A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for personalized recommendation of securities information based on knowledge graphs. Background Technology
[0002] In the process of recommending securities information, the recommendation method usually relies on users' historical clicks or simple tag matching. This method provides a coarse-grained portrayal of user interests and makes it difficult to dynamically reflect the characteristics of user interests changing over time. At the same time, it lacks in-depth exploration of the causal relationship between market events and their transmission impact on the industrial chain and individual stocks. As a result, the recommended content does not match the user's current focus and potential investment needs well, which affects the accuracy and personalization of the recommendation results. Summary of the Invention
[0003] This application provides a knowledge graph-based method and system for personalized securities information recommendation, which addresses the technical problems of insufficient accuracy and personalization in existing securities information recommendations.
[0004] In view of the above problems, this application provides a method and system for personalized recommendation of securities information based on knowledge graphs.
[0005] The first aspect of this application provides a personalized recommendation method for securities information based on knowledge graphs, the method comprising: Based on recommendation association factors of securities information, a multi-source knowledge graph is constructed, which includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. Responding to user clicks and close actions on information, the click frequency and close duration are extracted to construct a user interest decay field and a cognitive gap detection model, used to analyze the user's fine-grained interest intensity and unmet information needs for entities. Using the current moment as a trigger condition, recent or upcoming event nodes are extracted from the event graph, and causal chains are extended along causal edges to generate at least one causal inference path. The terminal events of the causal inference path are sequentially mapped to the industry chain graph and the securities graph to obtain a set of affected target stocks. Each stock in the target stock set is matched with the user interest decay field, and a recommendation score is calculated for each stock based on the causal strength of the causal inference path. A recommendation result is generated based on the recommendation score, including the causal inference path, the information set corresponding to each key event of the path, and at least one stock from the target stock set.
[0006] A second aspect of this application provides a personalized securities information recommendation system based on knowledge graphs, the system comprising: The graph construction module is used to construct a multi-source knowledge graph based on recommendation association factors of securities information. This multi-source knowledge graph includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. The model construction module is used to respond to user clicks and closes on information, extracting the user's click frequency and close duration, and constructing a user interest decay field and cognitive gap detection model to analyze the user's fine-grained interest intensity and unmet information needs for entities. The path generation module is used to extract recent or upcoming event nodes from the event graph, using the current moment as a trigger condition, and perform path generation along causal edges. The process involves several steps: First, a causal chain extension is performed to generate at least one causal inference path. The terminal events of this path are then sequentially mapped to an industry chain graph and a securities graph to obtain a set of affected target stocks. A recommendation score calculation module matches each stock in the target stock set with the user interest decay field and calculates a recommendation score for each stock based on the causal strength of the causal inference path. A recommendation result generation module generates a recommendation result based on the recommendation score. This recommendation result includes the causal inference path, the information set corresponding to each key event in the path, and at least one stock from the target stock set.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application constructs a multi-source knowledge graph based on recommendation association factors of securities information. The multi-source knowledge graph includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. Responding to user clicks and close actions on information, it extracts the user's click frequency and close duration, constructing a user interest decay field and a cognitive gap detection model to analyze the user's fine-grained interest intensity in entities and unmet information needs. Using the current moment as a trigger condition, it extracts recent or upcoming event nodes from the event graph, extends causal chains along causal edges to generate at least one causal inference path, and sequentially maps the terminal events of the causal inference path to the industry chain graph and the securities graph to obtain an affected set of target stocks. Each stock in the target stock set is matched with the user interest decay field, and a recommendation score is calculated for each stock based on the causal strength of the causal inference path. A recommendation result is generated based on the recommendation score, including the causal inference path, the information set corresponding to each key event of the path, and at least one stock from the target stock set. This invention addresses the technical problems of insufficient accuracy and personalization in existing securities information recommendations. By constructing a multi-source knowledge graph and combining it with a user interest decay field to match and recommend target stocks affected by events, it achieves the technical effect of improving the accuracy and personalization of securities information recommendations. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A schematic diagram of the personalized securities information recommendation method based on knowledge graph provided in this application embodiment; Figure 2 A schematic diagram of the structure of a knowledge graph-based personalized recommendation system for securities information provided in this application embodiment.
[0010] Figure labeling: Graph construction module 11, Model construction module 12, Path generation module 13, Recommendation score calculation module 14, Recommendation result generation module 15. Detailed Implementation
[0011] This application provides a knowledge graph-based personalized securities information recommendation method and system, addressing the technical problems of insufficient accuracy and personalization in existing securities information recommendations. By constructing a multi-source knowledge graph and combining it with user interest decay fields to match and recommend target stocks affected by events, the technical effect of improving the accuracy and personalization of securities information recommendations is achieved.
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0013] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0014] Example 1, as Figure 1 As shown, this application provides a personalized recommendation method for securities information based on knowledge graphs, the method comprising: Step S100: Based on the recommendation association factors of securities information, construct a multi-source knowledge graph, which includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph.
[0015] Furthermore, the method provided in the application embodiments also includes: The event graph uses events as nodes and causal and temporal relationships between events as edges; the industry chain graph uses companies, industries, and products as nodes and upstream and downstream supply and competition relationships as edges; the securities graph uses stocks, concepts, and financial indicators as nodes and attribution and correlation relationships as edges; the user behavior graph uses users, information, click behavior, and close behavior as nodes, and records the user's click timestamp and dwell time with click behavior edges and close behavior edges.
[0016] In this embodiment, a multi-source knowledge graph is constructed based on the recommendation association factors of securities information. Specifically, firstly, raw data related to securities information recommendations is acquired, including at least securities information data, industry data, securities data, and user behavior data. Then, the raw data is preprocessed, including standardizing names, codes, time formats, and identifiers, so that data representing the same object from different data sources can be mapped to the same entity. Next, graphs are constructed according to event information, industry information, securities information, and user behavior information respectively, resulting in an event graph, an industry chain graph, a securities graph, and a user behavior graph. Finally, the event graph, industry chain graph, securities graph, and user behavior graph are associated and organized to form a multi-source knowledge graph.
[0017] When constructing the event graph, event information is first extracted from securities data, and each event that can represent market changes, industry changes, or company changes is identified as an event node. Event nodes can record the event name, event type, event subject, event occurrence time, and event description. Then, event relationship identification is performed on any two event nodes, including causal relationship identification and temporal relationship identification. Specifically, the process first compares the occurrence times of the events corresponding to two event nodes. Only if the occurrence time of the event corresponding to the first event node is earlier than that of the event corresponding to the second event node is a further determination of whether a correlation exists between the two. Next, the process compares the event subjects, industry objects, product objects, or keyword information corresponding to the two event nodes. When a correspondence exists between the two in terms of subject, industry, product, or keyword, a basis for correlation is established. Then, the process identifies trigger words or expressions indicating causal meaning from the corresponding information content. Trigger words or expressions include terms such as "leads to," "initiates," "promotes," "drives," and "influences." When it is identified that the first event triggers, promotes, or influences the second event, a causal relationship is established between the first and second event nodes. Simultaneously, based on the chronological order of the events, a temporal relationship is established between event nodes that occurred earlier and those that occurred later. This results in an event graph with events as nodes and causal and temporal relationships between events as edges.
[0018] When constructing a supply chain map, the first step is to extract company, industry, and product information from industry data, and then identify these companies, industries, and products as nodes in the supply chain map. Next, the business relationships between companies, industries, and products are identified. Specifically, this involves retrieving information on the company's main business, products, suppliers, customers, and industry classifications. When a product, raw material, component, or service provided by Company One is used as a production input, operating input, or business dependency by Company Two, an upstream / downstream supply relationship is established between the corresponding nodes. Similarly, when a product belongs to a specific industry's production stage, or a company is positioned at a corresponding stage in an industry chain, an upstream / downstream supply relationship can also be established between the corresponding nodes. Furthermore, for identifying competitive relationships, the main businesses, product categories, and service targets of different companies are compared, or the uses, functions, and application scenarios of different products are compared. When two companies are in the same sub-industry and provide the same or similar products, or when two products are substitutable, a competitive relationship is established between the corresponding nodes. This results in a supply chain map with companies, industries, and products as nodes, and upstream / downstream supply relationships and competitive relationships as edges.
[0019] When constructing a security graph, the first step is to extract stock, concept, and financial indicator information from security data, and then define these as nodes in the security graph. Subsequently, the attribution relationships between stocks and concepts, as well as the correlation relationships between stocks and financial indicators, are identified. Specifically, the industry classification, theme classification, sector classification, or concept tag corresponding to a stock is read. When a stock is marked as belonging to a certain concept, an attribution relationship is established between the stock node and the corresponding concept node. Further, the financial data corresponding to the stock is read, extracting financial indicators such as operating revenue, net profit, gross profit margin, debt-to-equity ratio, and price-to-earnings ratio, and correlation relationships are established between the stock node and the corresponding financial indicator node. For cases where the same stock corresponds to multiple concepts or multiple financial indicators, multiple attribution and correlation relationships are established separately. Thus, a security graph is obtained with stocks, concepts, and financial indicators as nodes, and attribution and correlation relationships as edges.
[0020] When constructing a user behavior graph, the first step is to extract user interaction records with information from user behavior data, and then identify the user, information, click behavior, and close behavior from each interaction record. Specifically, when it is detected that a user opens a piece of information, the user identifier, information identifier, and access time corresponding to that access are extracted to generate a click behavior node; when it is detected that a user exits the information page, closes the information page, or switches to another page to end the current browsing, the corresponding user identifier, information identifier, and end time are extracted to generate a close behavior node. Then, the user node, information node, click behavior node, and close behavior node are written into the user behavior graph and connected through click behavior edges and close behavior edges. Specifically, when a click behavior occurs, a click timestamp is recorded; after a close behavior occurs, the time difference between the close time and the click time is calculated to obtain the dwell time, and the click timestamp and dwell time are recorded on the click behavior edge and close behavior edge. Thus, a user behavior graph is obtained with users, information, click behavior, and close behavior as nodes, and click behavior edges and close behavior edges recording the user's click timestamp and dwell time.
[0021] After obtaining the event graph, industry chain graph, securities graph, and user behavior graph, these four types of graphs are linked and organized to form a unified multi-source knowledge graph. Specifically, event information extracted from securities news is mapped to event nodes in the event graph; company, industry, and product information involved in the securities news is mapped to corresponding nodes in the industry chain graph; stock, concept, and financial indicator information involved in the securities news is mapped to corresponding nodes in the securities graph; and user click and close behaviors related to the securities news are mapped to corresponding nodes in the user behavior graph. Through this method, the same securities news can be simultaneously associated with event information, industry information, securities information, and user behavior information, thereby completing the construction of a multi-source knowledge graph based on recommendation association factors for securities news.
[0022] Step S200: In response to the user's click and close operations on the information, extract the user's click frequency and close duration, construct a user interest decay field and cognitive gap detection model, which is used to analyze the user's fine-grained interest intensity on entities and unmet information needs.
[0023] In this embodiment, in response to a user's click and close operations on information, when extracting the user's click frequency and close duration, and constructing a user interest decay field and cognitive gap detection model, the following steps are first taken: First, the unique identifier of the information, the click timestamp, the close timestamp, and the entity associated with the information in the knowledge graph corresponding to the user's click and close operations are recorded. The close duration is then obtained based on the click timestamp and close timestamp. Next, the historical interaction records of the same user are organized into an interaction sequence according to time order, and the instantaneous interest intensity value corresponding to each click is generated by combining the click frequency and close duration. Then, a user interest decay field is constructed based on the interaction sequence and instantaneous interest intensity value to analyze the user's fine-grained interest intensity on entities. Subsequently, a cognitive gap detection model is constructed based on the user's click frequency, average close duration, and information entropy increment of new information under the corresponding topic to analyze the user's unmet information needs.
[0024] Furthermore, in the method provided in the application embodiments, in response to the user's click and close operations on information, the user's click frequency and close duration are extracted to construct a user interest decay field and cognitive gap detection model, and the method further includes: In response to a user's click on any piece of information, the system records the unique identifier of the clicked information, the click timestamp, and at least one entity associated with the corresponding information in the knowledge graph. In response to a user's close operation on the same piece of information, the system records the close timestamp and calculates the dwell time based on the difference between the close timestamp and the click timestamp. All historical interaction records of the same user are organized into an interaction sequence in chronological order. For each click in the interaction sequence, based on the deviation between its dwell time and the preset expected dwell time, and combined with the logarithmic compression transformation of the dwell time, a learnable nonlinear mapping is used to generate the instantaneous interest intensity value corresponding to this click. Based on the interaction sequence and the instantaneous interest intensity value, a user interest decay field is constructed. The user interest decay field uses the user, entity, and time point as independent variables and outputs the user's cumulative interest intensity on a specified entity at the current moment. Furthermore, the system calculates the ratio of the user's click frequency to the average dwell time on the same topic of information and combines it with the information entropy increment of new information under the corresponding topic to quantify the user's cognitive gap on the topic and construct a cognitive gap detection model.
[0025] In this embodiment, in response to a user's click on any piece of information, the unique identifier of the clicked information, the click timestamp, and at least one entity associated with the corresponding information in the knowledge graph are recorded. Specifically, when a user clicks on a piece of information, the information number corresponding to that information is extracted as a unique identifier, the time of the click is extracted as a click timestamp, and the corresponding information node is located in the knowledge graph based on the unique identifier. Then, the entity nodes that are associated with the information node are read to obtain at least one entity associated with the corresponding information in the knowledge graph, wherein the entity is a company, industry, product, event, or stock in the knowledge graph. Thus, a click record containing the user identifier, unique identifier, click timestamp, and entity is formed.
[0026] Next, in response to the user's close action on this information, a close timestamp is recorded, and the dwell time is calculated based on the difference between the close timestamp and the click timestamp. Specifically, when the user ends browsing the current information, the moment of closing is extracted as the close timestamp, and this close timestamp is paired with the click timestamp in the click record to calculate the dwell time. The dwell time represents the user's continuous browsing time of this information. For records with abnormal dwell times, records with dwell times less than zero are removed, and dwell times exceeding a preset limit are limited to the preset limit to ensure that subsequent calculations are based on valid interaction records. Thus, a single interaction record containing a unique identifier, click timestamp, close timestamp, dwell time, and entity is obtained.
[0027] After obtaining multiple single-interaction records from the same user, all historical interaction records of the same user are organized into an interaction sequence in chronological order. Specifically, the earlier single-interaction records are arranged first, followed by the later single-interaction records, based on the click timestamp. When there are cases with the same click timestamp, they are then arranged in order of closing timestamp. Each record in the sorted interaction sequence retains a unique identifier, click timestamp, closing timestamp, dwell time, and physical entity corresponding to that interaction, thus obtaining an interaction sequence that reflects the historical interaction evolution process of the same user.
[0028] For each click in the interaction sequence, based on the deviation between its dwell time and the preset expected dwell time, and combined with a logarithmic compression transformation of the dwell time, a learnable nonlinear mapping is used to generate the instantaneous interest intensity value corresponding to this click. Specifically, firstly, the expected dwell time is determined according to the information type, length, or topic of the current information. The expected dwell time is used to characterize the reference time required to complete normal reading of the information. Then, the deviation between the dwell time and the expected dwell time is calculated, i.e. Where T represents the duration of stay. Indicates the expected length of stay. This indicates the deviation of the actual browsing level from the reference browsing level. Subsequently, a logarithmic compression transformation is applied to the dwell time. Here, L is used to reduce the amplifying effect of excessively long dwell times on the results while preserving the regularity of dwell time. Then, δ and L are input into a learnable nonlinear mapping to obtain the instantaneous interest intensity. , where S is the instantaneous interest intensity value, which takes a value between 0 and 1, and a, b and c are mapping parameters.
[0029] In this process, the parameters a, b, and c in the learnable nonlinear mapping are determined through learning from historical interaction samples. Specifically, training samples are first extracted from historical interaction records, and each training sample includes at least the dwell time T corresponding to that interaction and the expected dwell time. The parameters are: deviation degree δ, logarithmic compression result L, and target interest label y. The target interest label y represents the actual degree of interest corresponding to the click and can be determined based on the user's subsequent interactions with the same entity or topic after the click. If the user clicks on related information again within a preset time window and the subsequent dwell time is greater than or equal to the corresponding expected dwell time, the target interest label y is set to 1; if the user does not click on related information again within the preset time window, or if the user clicks on related information again but the subsequent dwell time is less than the corresponding expected dwell time, the target interest label y is set to 0. Then, δ and L from each training sample are substituted into a nonlinear mapping to obtain the predicted instantaneous interest intensity value S, and the difference between the predicted instantaneous interest intensity value S and the target interest label y is calculated. Next, the differences of all training samples are summarized to obtain the overall deviation corresponding to the current parameters a, b, and c; then, parameters a, b, and c are adjusted according to the overall deviation so that the adjusted parameters can reduce the difference between the predicted instantaneous interest intensity value S and the target interest label y. After each parameter adjustment, the updated parameters a, b, and c are used to recalculate the prediction results and corresponding differences for each training sample. This parameter adjustment is repeated until the overall deviation is less than a preset threshold or the preset number of iterations is reached. Through this learning process, parameters a, b, and c that match historical interaction samples are determined, enabling the instantaneous interest intensity value output by the learnable nonlinear mapping to reflect changes in the user's actual interest.
[0030] After obtaining the instantaneous interest intensity value corresponding to each click in the interaction sequence, a user interest decay field is constructed based on the interaction sequence and the instantaneous interest intensity values. Specifically, the click records in the interaction sequence are first grouped by entity, so that all instantaneous interest intensity values of the same user for the same entity are concentrated under the same entity item; then, when calculating the interest in a specified entity at the current moment, time decay is introduced for each historical click. For the i-th click, let its click time be . The current time is The time interval is Then, determine the attenuation coefficient corresponding to the click based on the time interval. in, The preset decay coefficient controls the rate at which interest decreases over time. A larger time interval results in a smaller decay coefficient, indicating that earlier clicks contribute less to the current interest; conversely, a smaller time interval results in a larger decay coefficient, indicating that more recent clicks contribute more to the current interest. Subsequently, the instantaneous interest intensity values of the same user for the same entity are multiplied by their corresponding decay coefficients and accumulated to obtain the user's cumulative interest intensity for the specified entity at the current moment. ,in, Indicates that user u is at the current time The cumulative intensity of interest in entity e, This represents the instantaneous interest intensity value corresponding to the i-th click. This represents the decay coefficient corresponding to the i-th click. Thus, a user interest decay field is constructed, which uses the user, entity, and time point as independent variables, and outputs the cumulative interest intensity of the user towards a specified entity at the current moment.
[0031] Finally, the ratio of user click frequency to average dwell time for the same topic is statistically analyzed, and combined with the information entropy increment of new information under the corresponding topic, the cognitive gap of users on this topic is quantified. Specifically, firstly, concept nodes or information clusters in the knowledge graph are defined as topics, and the total number of clicks and average dwell time for each topic within a preset time window are counted; then, the information entropy increment of each topic within the current time window is calculated to represent the average amount of new information carried by new information under the corresponding topic; based on this, the total number of clicks is divided by the sum of the average dwell time and a minimum constant, and then multiplied by the information entropy increment to obtain the cognitive gap index of the corresponding topic; finally, the cognitive gap index is compared with a preset threshold. When the cognitive gap index exceeds the preset threshold, it is determined that there is an unmet cognitive need for the corresponding topic, and the quantified cognitive gap index is output, thus obtaining the cognitive gap detection model.
[0032] Furthermore, the method provided in the application embodiments, which calculates the ratio of the user's click frequency to the average dwell time on the same topic, and combines this with the information entropy increment of new information under the corresponding topic to quantify the user's cognitive gap regarding the topic and construct a cognitive gap detection model, also includes: Concept nodes or information clusters in the knowledge graph are defined as topics. The total number of clicks and average dwell time for each topic within a preset time window are counted. The information entropy increment of each topic within the current time window is calculated to reflect the average amount of new information carried by newly published information in this topic. The total number of clicks is divided by the sum of the average dwell time and a minimum constant, and multiplied by the information entropy increment to obtain the cognitive gap index of the corresponding topic. When the cognitive gap index exceeds a preset threshold, it is determined that there is an unmet cognitive need for the corresponding topic, and a quantified cognitive gap index is output as the output of the cognitive gap detection model.
[0033] In this embodiment, when defining a concept node or information cluster in a knowledge graph as a topic, the concept nodes associated with securities information in the knowledge graph are first read, and each concept node is assigned to a topic. For information not directly associated with a concept node, keywords, entities, and event elements are extracted from the information. Based on keyword overlap, entity overlap, and event element overlap, information with similar content is grouped into information clusters, and each information cluster is assigned to a topic. Subsequently, a correspondence is established between each piece of information and its respective topic, and click records and dwell records generated by users for each topic are extracted within a preset time window. The number of click records is counted for each topic to obtain the total number of clicks for each topic. Then, the dwell time corresponding to each click under that topic is summed, and the summation result is divided by the total number of clicks to obtain the average dwell time for that topic. Through the above processing, the total number of clicks and average dwell time of users for each topic within the preset time window are obtained.
[0034] To calculate the information entropy increment of each topic within the current time window, firstly, newly published information belonging to the corresponding topic within the current time window is extracted. Keywords, entities, or event elements are then extracted from these newly published information, and the frequency of each keyword, entity, or event element in the newly published information for that topic is counted. Next, each frequency is divided by the total frequency to obtain the information distribution probability of each keyword, entity, or event element within the current time window. Then, based on the information distribution probability, the product of each probability value and its logarithm is calculated, and the sum of all products is taken as the negative to obtain the information entropy of that topic within the current time window. Based on this, historical information belonging to the same topic within the previous or baseline time window is extracted, and the frequency of keywords, entities, or event elements is counted in the same way to obtain the corresponding information distribution probability. The information entropy within the previous or baseline time window is then calculated based on this information distribution probability. Finally, the information entropy within the current time window is subtracted from the information entropy within the previous or baseline time window to obtain the information entropy increment of that topic within the current time window. Information entropy increment is used to characterize the degree of uncertainty change in newly added information on this topic within the current time window, thereby reflecting the average amount of new information carried by newly published information on this topic.
[0035] To obtain the cognitive gap index for a given topic, the process involves dividing the total number of clicks by the sum of the average dwell time and a minimum constant, and then multiplying this by the information entropy increment. First, the total number of clicks, average dwell time, and information entropy increment for that topic are read. Then, the total number of clicks is divided by the sum of the average dwell time and a minimum constant (a preset value used to avoid denominator anomalies when the average dwell time is zero or close to zero). Finally, the result is multiplied by the information entropy increment to obtain the cognitive gap index for that topic. The cognitive gap index quantifies the degree of cognitive gap a user has regarding a topic, allowing the total number of clicks, average dwell time, and information entropy increment to collectively represent the unmet cognitive needs of users.
[0036] Finally, the cognitive gap indicators for each topic are compared with preset thresholds. When the cognitive gap indicator for a topic is greater than the preset threshold, it is determined that there is an unmet cognitive need for that topic, and the corresponding cognitive gap indicator is output. When the cognitive gap indicator for a topic is less than or equal to the preset threshold, it is determined that there is no unmet cognitive need for that topic, or a corresponding low gap result is output. Subsequently, the determination results and quantified cognitive gap indicators for each topic are organized according to topic identifiers to form a cognitive gap detection model.
[0037] Step S300: Using the current moment as the trigger condition, extract recent or upcoming event nodes from the event graph, extend the causal chain along the causal edge to generate at least one causal deduction path, and map the terminal events of the causal deduction path to the industry chain graph and the securities graph in sequence to obtain the set of affected target stocks.
[0038] In this embodiment, when extracting recent or upcoming event nodes from the event graph using the current time as the trigger condition and extending the causal chain along the causal edges, firstly, using the current time as the time reference point, historical event nodes located within a first preset time window and predicted event nodes whose probability of occurrence exceeds a preset probability threshold within a second preset time window are selected from the event graph. These historical and predicted event nodes together constitute a seed event node set. Then, causal strength weights are associated with the directed causal edges in the event graph, and multi-hop extension is performed along the directed direction of the causal edges based on each node in the seed event node set. During the expansion process, the expansion direction is filtered based on the causal strength weight, and the cumulative causal strength of the path is recorded. Further, when the expansion depth reaches the preset maximum inference depth, or when there is no outgoing edge that meets the conditions at the current node, the expansion of the corresponding path is terminated, and the sequence of nodes and directed edges traversed from the starting node to the ending node are taken as candidate causal inference paths. Finally, the candidate causal inference paths are sorted according to the cumulative causal strength, and at least one of the top-ranked paths is selected as the final causal inference path output. The output results include the causal inference path, the cumulative causal strength value, and the sequence of event nodes involved in the path.
[0039] Next, the terminal events of the causal deduction path are sequentially mapped to the industry chain graph and the securities graph. In this process, firstly, the terminal event node of at least one causal deduction path is obtained; the terminal event node is a specific market event located at the end of the causal chain in the event graph. Then, through a pre-constructed set of event-industry chain mapping edges, the terminal event node is mapped to one or more industry nodes or product nodes in the industry chain graph affected by this event, obtaining a first intermediate mapping set. Subsequently, through a pre-constructed set of industry chain-securities mapping edges, the industry nodes or product nodes in the first intermediate mapping set are mapped to one or more corresponding stock nodes in the securities graph, obtaining a second intermediate mapping set. Finally, all stock nodes in the second intermediate mapping set are deduplicated and merged to obtain the set of affected target stocks.
[0040] Furthermore, in the method provided in the application embodiment, taking the current moment as the trigger condition, extracting recent or upcoming event nodes from the event graph, extending the causal chain along the causal edge, and generating at least one causal inference path, it also includes: Using the current moment as the time reference point, historical event nodes whose timestamps occur within a first preset time window before the current moment are selected from the event graph, along with predicted event nodes whose probability of occurrence within a second preset time window exceeds a preset probability threshold, as determined by a time series prediction model. These historical and predicted event nodes are collectively used as a seed event node set. A causal strength weight is pre-associated with each directed causal edge in the event graph, where the causal strength weight is determined based on historical time series data. Starting from each node in the seed event node set, a heuristic graph search algorithm is used to perform multi-hop expansion along the directed direction of the causal edge. In each hop expansion... Only outgoing edges with causal strength weights greater than a preset edge weight threshold are traversed, and the cumulative causal strength from the starting node to the current node is recorded. When the expansion depth reaches the preset maximum inference depth, or when the current node does not have an outgoing edge that satisfies the edge weight threshold, the expansion of this path is terminated, the current node is taken as the end node, and the sequence of nodes and directed edges traversed from the starting node to the end node is recorded as a candidate causal inference path. All candidate causal inference paths are sorted from high to low according to their cumulative causal strength, and at least one of the top-ranked paths is selected as the final causal inference path output. Each path is accompanied by its cumulative causal strength value and the sequence of event nodes involved in the path.
[0041] In this embodiment, when the current moment is taken as the time reference point, the occurrence timestamps corresponding to all event nodes in the event graph are first read, and the event nodes within the first preset time window before the current moment are filtered as historical event nodes. Simultaneously, for candidate events that have not yet occurred but may occur in the future, a time series prediction model for outputting the probability of event occurrence is first constructed. Specifically, the occurrence records of various events in continuous time intervals are extracted from historical time series data, forming an event occurrence sequence according to time order. The occurrence frequency, occurrence interval, co-occurrence, and event type distribution of each event within a preset length of historical time slice are used as input features. These input features constitute the input data of the time series prediction model, and the input data is arranged in time slice order to form a time series input sequence. Whether the corresponding event occurs within the subsequent second preset time window is used as the output label. The output label constitutes the output data of the time series prediction model, and the output data is used to characterize the occurrence result of each candidate event within the future second preset time window, thus training the time series prediction model. During model training, the model parameters of the time series prediction model include at least an input weight matrix, a state transition weight matrix, a bias vector, and an output weight matrix. The input weight matrix is used to represent... The model assesses the impact of input data on the current time series state. The state transition weight matrix represents the influence of the previous time series state on the current time series state. The bias vector adjusts the state update result by offsetting the bias. The output weight matrix maps the time series state to the probability of occurrence of the corresponding event. During training, the input data is fed into the time series prediction model to obtain the model output. The model output and the output data are then compared to calculate the error. Based on the error, the input weight matrix, state transition weight matrix, bias vector, and output weight matrix are iteratively updated until a preset training termination condition is met. After training, the event sequences from multiple consecutive time slices before the current time are input into the time series prediction model to obtain the probability of occurrence of each candidate event within a second preset time window. Candidate events with a probability exceeding a preset probability threshold are identified as predicted event nodes. Thus, historical event nodes and predicted event nodes are uniformly included in a seed event node set. Historical event nodes represent market events that have occurred near the current time, while predicted event nodes represent events that meet probability conditions within a future time range. The seed event node set serves as the source of starting nodes for subsequent causal chain expansion.
[0042] After forming the seed event node set, the directed causal edges between event nodes in the event graph are read, and a causal strength weight is pre-associated for each directed causal edge. Specifically, the starting event of a directed causal edge is recorded as the preceding event, and the ending event as the following event. Then, the total number of occurrences of the preceding event is counted in the historical time series data, and the co-occurrence count of the following event following the preceding event within a preset propagation time window is counted. The co-occurrence count is then divided by the total number of occurrences of the preceding event to obtain the conditional occurrence ratio of the preceding event to the following event. Next, the average time interval between the occurrence of the following event after the preceding event is counted, and the corresponding preset time coefficient is retrieved based on the average time interval. The conditional occurrence ratio is multiplied by the time coefficient to obtain the causal strength weight corresponding to the directed causal edge. Thus, each directed causal edge in the event graph has a corresponding causal strength weight value, providing a unified basis for edge selection and path comparison in subsequent path expansion.
[0043] Then, each node in the seed event node set is used as the starting node in turn, and multi-hop expansion is performed along the directed direction of the causal edges. At the start of the expansion, the current path is initialized to contain only the starting node, and the initial cumulative causal strength of the path is set to a preset initial value. Then, all outgoing edges of the current node are read, and the causal strength weight of each outgoing edge is compared with the preset edge weight threshold. Only outgoing edges with a causal strength weight greater than the preset edge weight threshold are retained as valid expansion outgoing edges. The next event node is then visited along the valid expansion outgoing edges, and the next event node is added to the node sequence of the current path. At the same time, the corresponding directed causal edge is added to the directed edge sequence of the current path. After each hop expansion, the cumulative causal strength from the starting node to the current node is updated. The cumulative causal strength is obtained by multiplying the causal strength weights of each directed causal edge in the current path in the order of the path. During the expansion process, the heuristic graph search algorithm uses the cumulative causal strength of the current path as the path priority, and prioritizes the expansion of paths with larger cumulative causal strength values, thereby advancing the causal chain expansion along the direction of stronger causal connections.
[0044] As multi-hop expansion continues, the number of hops traversed, node sequence, directed edge sequence, and corresponding cumulative causal strength of the current path are recorded synchronously. When the expansion depth reaches the preset maximum inference depth, the expansion of this path is stopped, or the expansion of this path is terminated when the current node has no outgoing edges with a causal strength weight greater than a preset edge weight threshold. After the expansion terminates, the current node is determined as the end node, and all event nodes traversed from the start node to the end node are recorded as a node sequence in the order of access. All directed causal edges connecting adjacent event nodes are recorded as a directed edge sequence in the order of access. Combined with the cumulative causal strength of the path, a candidate causal inference path is formed. The end node represents the termination event of the current path in the causal chain expansion, and the candidate causal inference path represents a complete event chain obtained by hop-by-hop inference from the seed event node along the directed causal edges.
[0045] After obtaining all candidate causal inference paths, the cumulative causal strength value corresponding to each candidate causal inference path is read and sorted from high to low according to the cumulative causal strength value. Then, according to the preset output quantity, at least one candidate causal inference path with the highest ranking is selected from the sorted results as the final causal inference path output. During output, the cumulative causal strength value, the sequence of event nodes involved in the path, and the corresponding directed edge sequence for each final causal inference path are output together. The event node sequence represents all event nodes covered by the final causal inference path, the directed edge sequence represents the causal connection relationship between adjacent event nodes, and the cumulative causal strength value represents the overall path weight corresponding to the final causal inference path. Thus, the entire process of extracting recent or upcoming event nodes from the event graph, extending the causal chain along causal edges, and generating at least one causal inference path is completed, using the current moment as the trigger condition.
[0046] Furthermore, the method provided in the application embodiments, which sequentially maps the terminal events of the causal deduction path to an industry chain graph and a security graph to obtain a set of affected target stocks, also includes: Obtain the terminal event node of at least one causal deduction path, where the terminal event node is a specific market event located at the end of the causal chain in the event graph; map the terminal event node to one or more industry nodes or product nodes in the industry chain graph affected by the event through a pre-constructed event-industry chain mapping edge set to obtain a first intermediate mapping set; map each industry node or product node in the first intermediate mapping set to one or more corresponding stock nodes in the securities graph through a pre-constructed industry chain-securities mapping edge set to obtain a second intermediate mapping set; deduplicate and merge all stock nodes in the second intermediate mapping set to obtain the affected target stock set.
[0047] In this embodiment, after obtaining at least one causal deduction path, the causal deduction path is first sequentially parsed, and the last event node is extracted from the event node sequence corresponding to the causal deduction path, which is then identified as the terminal event node. The terminal event node is a specific market event located at the end of the causal chain in the event graph, used to characterize the triggering event that ultimately falls at the industry impact level after being propagated step by step from preceding events. After extraction, the event type, event subject, event object, and event impact label corresponding to the terminal event node are read to provide the terminal event node with the input information required to perform mapping to the industry chain graph.
[0048] After identifying the terminal event node, a pre-constructed set of event-industry chain mapping edges is used to map the terminal event node to one or more industry nodes or product nodes in the industry chain graph that are affected by the event. Specifically, the event-industry chain mapping edge set pre-establishes the correspondence between event nodes in the event graph and industry nodes and product nodes in the industry chain graph, and records the influence coefficient on each event-industry chain mapping edge. The influence coefficient is obtained by regression calculation using historical industry index change data after the event occurs, and is used to characterize the influence weight when the terminal event node transmits to industry nodes or product nodes. When performing the mapping, the corresponding event-industry chain mapping edge is first searched in the event-industry chain mapping edge set according to the event type of the terminal event node. Then, the industry nodes or product nodes connected to the event-industry chain mapping edge are read, and industry nodes or product nodes with influence coefficients greater than a preset mapping threshold are included in the mapping result. After the above processing, a first intermediate mapping set is obtained by mapping the terminal event node. The first intermediate mapping set is used to characterize the set of industry chain nodes affected by the terminal event node.
[0049] Subsequently, using a pre-constructed industry chain-securities mapping edge set, each industry node or product node in the first intermediate mapping set is mapped to one or more corresponding stock nodes in the securities graph. Specifically, the industry chain-securities mapping edge set pre-establishes the connection relationships between industry nodes, product nodes in the industry chain graph, and stock nodes in the securities graph, and records the edge weight on each industry chain-securities mapping edge to represent the association weight when the industry chain node is transmitted to the stock node. During the mapping process, each industry node or product node in the first intermediate mapping set is read sequentially, and then the stock node connected to it is searched in the industry chain-securities mapping edge set according to the node identifier of the industry node or product node. Stock nodes with edge weights greater than a preset securities mapping threshold are included in the mapping result. Then, combined with the status information of the corresponding stock nodes in the securities graph, the mapped stock nodes are filtered to remove stock nodes that are suspended from trading on the same day, stock nodes that are locked at the daily price limit, and stock nodes with a market capitalization lower than a preset threshold. After the above processing, a second intermediate mapping set is obtained by further mapping from the first intermediate mapping set. The second intermediate mapping set is used to represent the set of candidate stock nodes obtained after the transmission from the industry chain graph to the securities graph.
[0050] Finally, all stock nodes in the second intermediate mapping set are deduplicated and merged. This is done by reading the stock code corresponding to each stock node in the second intermediate mapping set one by one, and identifying stock nodes with the same stock code as duplicate nodes. Only one of these duplicate nodes is kept, and the remaining duplicate records are deleted. After cleaning up all duplicate nodes, the remaining stock nodes are aggregated to form the affected target stock set.
[0051] Furthermore, in the method provided in the application embodiments, the pre-constructed event-industry chain mapping edge set also includes: For each event type node in the event graph, the affected industry chain links are predefined, and a directed mapping edge is established from the event node to the industry node or product node in the industry chain graph. Each mapping edge is accompanied by an influence coefficient, which represents the transmission strength of the event to the target industry or product. The mapping edge and its influence coefficient are obtained by a combination of one or more methods, such as regression analysis based on industry index fluctuations after historical events, domain expert knowledge annotation, or automatic extraction of event-industry co-occurrence patterns from research reports and news using natural language processing technology to quantify the correlation strength.
[0052] In this embodiment, for each event type node in the event graph, the corresponding industry chain link it affects is predefined. Then, based on the industry chain link, the corresponding industry node or product node is located in the industry chain graph, and a directed mapping edge from the event node to the industry node or product node is established. If regression analysis based on the fluctuation of industry index after historical events is used to obtain the mapping edge and its influence coefficient, the construction method of the industry index is first determined. For a certain industry node, all stock nodes belonging to that industry node in the securities graph are read, the closing price of each stock node on the same trading day is extracted, and the rise and fall of each stock node relative to the previous trading day is calculated. Then, the rise and fall of all stock nodes are weighted and summed according to the circulating market value to obtain the industry index rise and fall sequence of that industry node on that trading day. For product nodes, the industry node to which the product node belongs is first determined, and then the industry index rise and fall sequence corresponding to that industry node is used. Subsequently, all historical occurrence dates of the event type node are extracted, and for each historical occurrence date, the cumulative increase or decrease of the industry index within a preset observation window after the event occurs is calculated. Simultaneously, control dates where the event did not occur are selected, and the cumulative increase or decrease of the industry index within the same observation window is calculated. Then, using whether the event occurred as input and the cumulative increase or decrease of the industry index as output, regression analysis is performed to obtain the regression coefficients of the event type node pointing to the corresponding industry node or product node. All regression coefficients are then normalized to a range of zero to one. This is done by first reading the maximum and minimum values of all regression coefficients, then subtracting the minimum value from the current regression coefficient, and dividing by the difference between the maximum and minimum values to obtain the normalized result. This normalized result is used as the influence coefficient of the corresponding directed mapping edge. If the influence coefficient is greater than a preset threshold, the directed mapping edge is retained; otherwise, it is deleted.
[0053] If the mapping edges and their influence coefficients are obtained by using domain expert knowledge annotation, then domain experts first annotate the corresponding industry chain links affected by each event type node, and further annotate the corresponding industry nodes or product nodes. Simultaneously, a score is given for each mapping relationship from an event node to an industry node or product node. The score can be a numerical label from zero to one, where zero represents no influence and one represents the strongest influence. The score is directly used as the influence coefficient. If the event-industry co-occurrence pattern is automatically extracted from research reports and news using natural language processing technology to quantify its correlation strength, then research report and news texts containing the event type node are first collected. Event words, industry names, and product names are extracted from the text, and the co-occurrence frequency of event words and industry names or product names within the same text segment is counted. The co-occurrence frequency is then divided by the total occurrence frequency of the event word to obtain the text association value from the event node to the corresponding industry node or product node. Afterward, all text association values are normalized in the same way as described above to obtain the influence coefficient, and directed mapping edges are established for mapping relationships with influence coefficients greater than a preset threshold. If the mapping edge and its influence coefficient are obtained by combining multiple methods, the influence coefficients corresponding to the regression analysis results, knowledge annotation results, and text extraction results are first obtained separately and then uniformly converted to the range of zero to one. The final influence coefficient is then calculated by weighted summation. The weights corresponding to each method are preset, and the sum of the weights is one. The weights can be determined based on the accuracy ratio of each method in historical samples. For example, the higher the accuracy of the regression analysis results, the greater its corresponding weight. The knowledge annotation results and text extraction results are also determined according to the same rules. Finally, the weighted summation result is used as the influence coefficient of the directed mapping edge, and the corresponding mapping edge is retained accordingly.
[0054] Furthermore, in the method provided in the application embodiments, both obtaining the first intermediate mapping set and obtaining the second intermediate mapping set adopt a cross-graph linkage query mechanism, which further includes: Event nodes, industry chain nodes, and security nodes are treated as different types of vertices in a heterogeneous graph. A predefined cross-graph element path template is used, which sequentially includes event node type, industry chain node type, and security node type. For the terminal event node of each causal inference path, a meta-path instantiation query is performed on the heterogeneous graph to return all industry chain nodes and security nodes that satisfy the meta-path template. The query process is based on the influence coefficient of the event-industry chain mapping edge and the weight of the industry chain-security mapping edge. Only when both are greater than their respective preset thresholds are the corresponding security nodes included in the target stock set. The inclusion in the target stock set also includes excluding stocks that are suspended from trading on the same day, locked at the daily price limit, or have a market capitalization lower than a preset threshold, based on the stock fundamental indicators and real-time market data pre-stored in the security graph.
[0055] In this embodiment, when representing event nodes, industry chain nodes, and securities nodes uniformly in the same heterogeneous graph, event nodes are first read from the event graph, industry nodes and product nodes are read from the industry chain graph, and stock nodes are read from the securities graph, each assigned a different node type identifier. The event node type represents market events in the event graph, the industry chain node type represents industry nodes and product nodes in the industry chain graph, and the securities node type represents stock nodes in the securities graph. Subsequently, event-industry chain mapping edges and industry chain-securities mapping edges are written into the heterogeneous graph, enabling event nodes, industry chain nodes, and securities nodes to form a unified graph structure through cross-graph connections. After completing the heterogeneous graph construction, a predefined cross-graph element path template is defined, which sequentially includes event node type, industry chain node type, and securities node type. This specifies that the query path starts from the event node, first reaches the industry chain node, and then reaches the securities node. Thus, the cross-graph element path template is fixed as a query structure of event node type, industry chain node type, and securities node type, constraining the node access order for subsequent cross-graph queries.
[0056] For each terminal event node in a causal deduction path, when performing a meta-path instantiation query on the heterogeneous graph, the terminal event node corresponding to the causal deduction path is first read and used as the starting node for the meta-path instantiation query. Then, according to a predefined cross-graph meta-path template, starting from the terminal event node in the heterogeneous graph, the supply chain nodes directly connected to the terminal event node are retrieved. From the retrieved supply chain nodes, the directly connected securities nodes are then retrieved, thus obtaining a node combination that satisfies the path structure of event node type, supply chain node type, and securities node type. The meta-path instantiation query is used to find actual node paths in the heterogeneous graph that conform to the predefined node type order. The returned results include supply chain nodes connected to the terminal event node and securities nodes obtained by further connecting supply chain nodes. When executing a query, the influence coefficient of the event-industry chain mapping edge between the event node and the industry chain node is read, and the influence coefficient is compared with the corresponding preset threshold. Only industry chain nodes with influence coefficients greater than the corresponding preset threshold are retained for subsequent queries. Then, the edge weight of the industry chain-securities mapping edge between the retained industry chain node and the securities node is read, and the edge weight is compared with the corresponding preset threshold. Only securities nodes with edge weights greater than the corresponding preset threshold are retained as valid query results.
[0057] After obtaining security nodes that meet the meta-path template and pass the threshold screening, the validity of these nodes is further filtered based on pre-stored stock fundamental indicators and real-time market data in the security graph. Specifically, the stock fundamental indicators and real-time market data corresponding to each security node are first read. The stock fundamental indicators include market capitalization, and the real-time market data includes the trading status and daily price limits. Then, it is checked whether the stock corresponding to the security node is suspended from trading on the current trading day. If it is suspended, the security node is removed. For security nodes that are not removed, it is further checked whether they are locked at the daily price limit. If they are locked at the daily price limit, the security node is removed. For the remaining security nodes, their market capitalization is read and compared with a preset threshold. If the market capitalization is lower than the preset threshold, the security node is removed. After the above screening, the remaining security nodes are determined as valid security nodes that can be included in the target stock set.
[0058] After the validity screening is completed, all retained securities nodes are aggregated according to their stock codes, and duplicate securities nodes with the same stock code are deduplicated and merged to form the target stock set. All stocks in the target stock set correspond to node paths that satisfy the cross-graphite path template, and the influence coefficient of the event-industry chain mapping edge in their path is greater than the corresponding preset threshold, the edge weight of the industry chain-securities mapping edge is greater than the corresponding preset threshold, and the stock is not suspended from trading, not under price limit lock-up, and its market capitalization is not lower than the preset threshold.
[0059] Step S400: Match each stock in the target stock set with the user interest decay field, and calculate the recommendation score for each stock by combining the causal strength of the causal inference path.
[0060] In this embodiment of the application, when matching each stock in the target stock set with the user interest decay field and combining the causal strength of the causal inference path to calculate the recommendation score of each stock, firstly, for each stock in the target stock set, the user's interest intensity value for the stock at the current moment is queried from the user interest decay field, and the corresponding cumulative causal strength value is obtained from the causal inference path that generated the stock; then, the interest intensity value and the cumulative causal strength value are fused by weighted summation to obtain the recommendation score of each stock.
[0061] Furthermore, in the method provided in the application embodiment, matching each stock in the target stock set with the user interest decay field, and calculating a recommendation score for each stock based on the causal strength of the causal inference path, further includes: For each stock in the target stock set, the user's interest intensity value for the stock at the current moment is queried from the user interest decay field, and the cumulative causal intensity value is obtained from the causal inference path that generated the stock. The interest intensity value and the cumulative causal intensity value are fused using a weighted summation method to obtain the recommendation score for each stock. The weight coefficient of the interest intensity value is a preset balance factor, which is used to control the relative importance of user personalized interests and event causal inference in the recommendation decision. The balance factor is dynamically adjusted according to the user's historical behavior or the current market volatility state. When the number of historical interactions of the user is lower than a preset threshold, the balance factor is reduced to enhance the contribution of causal inference; when the market volatility index exceeds a preset volatility threshold, the balance factor is increased to enhance the user's risk aversion interest selection.
[0062] In this embodiment, for each stock in the target stock set, an entity correspondence between the stock and the user interest decay field is first established. Specifically, the stock codes in the target stock set are read one by one, and the corresponding stock entity is searched in the user interest decay field based on the stock code. When a stock entity with the same stock code exists in the user interest decay field, the user's cumulative interest intensity for that stock entity at the current moment is directly read, and this cumulative interest intensity is determined as the user's interest intensity value for that stock at the current moment. The user interest decay field takes the user, entity, and time point as input and outputs the user's cumulative interest intensity for the specified entity at the current moment. Therefore, when performing a query, the current user identifier, the stock entity identifier corresponding to the current stock, and the current moment are all input into the user interest decay field to obtain the corresponding interest intensity value. After completing the interest intensity value query, the causal inference path of the original stock is read, and the corresponding cumulative causal intensity value is extracted from the causal inference path. When the same stock corresponds to multiple causal inference paths, the cumulative causal intensity values corresponding to each causal inference path are compared first. Then, the causal inference path with the largest cumulative causal intensity value is determined as the causal inference path of the original stock, and the cumulative causal intensity value corresponding to this path is used as the cumulative causal intensity value of this stock. In this way, each stock in the target stock set corresponds to an interest intensity value and a cumulative causal intensity value. The interest intensity value is used to represent the user's personalized interest level of attention to the stock at the current moment, and the cumulative causal intensity value is used to represent the path weight corresponding to the event causal inference link when it is transmitted to the stock.
[0063] After obtaining the interest intensity value and cumulative causal intensity value, a weighted summation method is used to merge the interest intensity value and the cumulative causal intensity value to obtain the recommendation score for each stock. Specifically, a preset balance factor is first assigned to the interest intensity value, denoted as α, and then 1 minus α is determined as the weight coefficient of the cumulative causal intensity value. Subsequently, the recommendation score is calculated as α multiplied by the interest intensity value, plus 1 minus α multiplied by the cumulative causal intensity value, to obtain the recommendation score for this stock. The preset balance factor is used to control the relative importance of user personalized interests and event causal inference in the recommendation decision. Therefore, before calculating the recommendation score, the preset balance factor is dynamically adjusted based on the user's historical behavior and the current market volatility. Specifically, during the adjustment, the number of historical interactions of the current user within a preset historical period is first read and compared with a preset threshold. When the number of historical interactions is lower than the preset threshold, the preset balance factor is reduced by a preset downward adjustment step size, so that the proportion of the interest intensity value in the recommendation score decreases, while the proportion of the cumulative causal intensity value in the recommendation score increases. After adjusting for the historical interaction counts, the current market volatility index is read and compared with a preset volatility threshold. When the market volatility index exceeds the preset threshold, the preset balance factor is increased by a preset step size based on the previous adjustment, increasing the proportion of interest intensity in the recommendation score to reflect the impact of user risk aversion choices on recommendation decisions in a high-volatility market. After dynamic adjustment, the adjusted preset balance factor is substituted into the weighted summation formula to obtain the recommendation score for each stock in the target stock set.
[0064] Step S500: Generate recommendation results based on the recommendation score. The recommendation results include the causal deduction path, the information set corresponding to each key event of the path, and at least one stock in the target stock set.
[0065] In this embodiment, when generating a recommendation result based on the recommendation score, the recommendation scores of each stock in the target stock set are first sorted, and at least one stock with the highest recommendation score is selected as the recommended stock. Then, for each recommended stock, the causal inference path that generated the stock is retrieved, the event nodes that pass through the causal inference path are extracted, and the event nodes are identified as key events of the path. Then, based on the event node identifiers corresponding to each key event, the associated information nodes are retrieved in the knowledge graph to obtain the information set corresponding to each key event of the path. Finally, the recommended stock, the causal inference path that generated the recommended stock, and the information set corresponding to each key event of the path are associated and organized to form the recommendation result, so that the recommendation result includes the causal inference path, the information set corresponding to each key event of the path, and at least one stock in the target stock set.
[0066] In summary, the embodiments of this application have at least the following technical effects: This application constructs a multi-source knowledge graph based on recommendation association factors of securities information. The multi-source knowledge graph includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. Responding to user clicks and close actions on information, it extracts the user's click frequency and close duration, constructing a user interest decay field and a cognitive gap detection model to analyze the user's fine-grained interest intensity in entities and unmet information needs. Using the current moment as a trigger condition, it extracts recent or upcoming event nodes from the event graph, extends causal chains along causal edges to generate at least one causal inference path, and sequentially maps the terminal events of the causal inference path to the industry chain graph and the securities graph to obtain an affected set of target stocks. Each stock in the target stock set is matched with the user interest decay field, and a recommendation score is calculated for each stock based on the causal strength of the causal inference path. A recommendation result is generated based on the recommendation score, including the causal inference path, the information set corresponding to each key event of the path, and at least one stock from the target stock set. This invention addresses the technical problems of insufficient accuracy and personalization in existing securities information recommendations. By constructing a multi-source knowledge graph and combining it with a user interest decay field to match and recommend target stocks affected by events, it achieves the technical effect of improving the accuracy and personalization of securities information recommendations.
[0067] Example 2, based on the same inventive concept as the knowledge graph-based personalized securities information recommendation method in the aforementioned examples, such as... Figure 2 As shown, this application provides a personalized securities information recommendation system based on knowledge graphs. The system and method embodiments in this application are based on the same inventive concept. The system includes: The graph construction module 11 is used to construct a multi-source knowledge graph based on the recommendation association factors of securities information. The multi-source knowledge graph includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. The model construction module 12 is used to respond to user clicks and closes on information, extract the user's click frequency and close duration, and construct a user interest decay field and cognitive gap detection model to analyze the user's fine-grained interest intensity and unmet information needs for entities. The path generation module 13 is used to extract recent or upcoming event nodes from the event graph, using the current moment as a trigger condition, and generate paths along causal edges. The causal chain is extended to generate at least one causal inference path, and the terminal events of the causal inference path are sequentially mapped to the industry chain graph and the securities graph to obtain the set of affected target stocks; the recommendation score calculation module 14 is used to match each stock in the target stock set with the user interest decay field, and calculate the recommendation score of each stock in combination with the causal strength of the causal inference path; the recommendation result generation module 15 is used to generate recommendation results based on the recommendation scores, and the recommendation results include the causal inference path, the information set corresponding to each key event of the path, and at least one stock in the target stock set.
[0068] Furthermore, the system is also used to implement the following functions: The event graph uses events as nodes and causal and temporal relationships between events as edges; the industry chain graph uses companies, industries, and products as nodes and upstream and downstream supply and competition relationships as edges; the securities graph uses stocks, concepts, and financial indicators as nodes and attribution and correlation relationships as edges; the user behavior graph uses users, information, click behavior, and close behavior as nodes, and records the user's click timestamp and dwell time with click behavior edges and close behavior edges.
[0069] Furthermore, the system is also used to implement the following functions: In response to a user's click on any piece of information, the system records the unique identifier of the clicked information, the click timestamp, and at least one entity associated with the corresponding information in the knowledge graph. In response to a user's close operation on the same piece of information, the system records the close timestamp and calculates the dwell time based on the difference between the close timestamp and the click timestamp. All historical interaction records of the same user are organized into an interaction sequence in chronological order. For each click in the interaction sequence, based on the deviation between its dwell time and the preset expected dwell time, and combined with the logarithmic compression transformation of the dwell time, a learnable nonlinear mapping is used to generate the instantaneous interest intensity value corresponding to this click. Based on the interaction sequence and the instantaneous interest intensity value, a user interest decay field is constructed. The user interest decay field uses the user, entity, and time point as independent variables and outputs the user's cumulative interest intensity on a specified entity at the current moment. Furthermore, the system calculates the ratio of the user's click frequency to the average dwell time on the same topic of information and combines it with the information entropy increment of new information under the corresponding topic to quantify the user's cognitive gap on the topic and construct a cognitive gap detection model.
[0070] Furthermore, the system is also used to implement the following functions: Concept nodes or information clusters in the knowledge graph are defined as topics. The total number of clicks and average dwell time for each topic within a preset time window are counted. The information entropy increment of each topic within the current time window is calculated to reflect the average amount of new information carried by newly published information in this topic. The total number of clicks is divided by the sum of the average dwell time and a minimum constant, and multiplied by the information entropy increment to obtain the cognitive gap index of the corresponding topic. When the cognitive gap index exceeds a preset threshold, it is determined that there is an unmet cognitive need for the corresponding topic, and a quantified cognitive gap index is output as the output of the cognitive gap detection model.
[0071] Furthermore, the system is also used to implement the following functions: Using the current moment as the time reference point, historical event nodes whose timestamps occur within a first preset time window before the current moment are selected from the event graph, along with predicted event nodes whose probability of occurrence within a second preset time window exceeds a preset probability threshold, as determined by a time series prediction model. These historical and predicted event nodes are collectively used as a seed event node set. A causal strength weight is pre-associated with each directed causal edge in the event graph, where the causal strength weight is determined based on historical time series data. Starting from each node in the seed event node set, a heuristic graph search algorithm is used to perform multi-hop expansion along the directed direction of the causal edge. In each hop expansion... Only outgoing edges with causal strength weights greater than a preset edge weight threshold are traversed, and the cumulative causal strength from the starting node to the current node is recorded. When the expansion depth reaches the preset maximum inference depth, or when the current node does not have an outgoing edge that satisfies the edge weight threshold, the expansion of this path is terminated, the current node is taken as the end node, and the sequence of nodes and directed edges traversed from the starting node to the end node is recorded as a candidate causal inference path. All candidate causal inference paths are sorted from high to low according to their cumulative causal strength, and at least one of the top-ranked paths is selected as the final causal inference path output. Each path is accompanied by its cumulative causal strength value and the sequence of event nodes involved in the path.
[0072] Furthermore, the system is also used to implement the following functions: Obtain the terminal event node of at least one causal deduction path, where the terminal event node is a specific market event located at the end of the causal chain in the event graph; map the terminal event node to one or more industry nodes or product nodes in the industry chain graph affected by the event through a pre-constructed event-industry chain mapping edge set to obtain a first intermediate mapping set; map each industry node or product node in the first intermediate mapping set to one or more corresponding stock nodes in the securities graph through a pre-constructed industry chain-securities mapping edge set to obtain a second intermediate mapping set; deduplicate and merge all stock nodes in the second intermediate mapping set to obtain the affected target stock set.
[0073] Furthermore, the system is also used to implement the following functions: For each event type node in the event graph, the affected industry chain links are predefined, and a directed mapping edge is established from the event node to the industry node or product node in the industry chain graph. Each mapping edge is accompanied by an influence coefficient, which represents the transmission strength of the event to the target industry or product. The mapping edge and its influence coefficient are obtained by a combination of one or more methods, such as regression analysis based on industry index fluctuations after historical events, domain expert knowledge annotation, or automatic extraction of event-industry co-occurrence patterns from research reports and news using natural language processing technology to quantify the correlation strength.
[0074] Furthermore, the system is also used to implement the following functions: Event nodes, industry chain nodes, and security nodes are treated as different types of vertices in a heterogeneous graph. A predefined cross-graph element path template is used, which sequentially includes event node type, industry chain node type, and security node type. For the terminal event node of each causal inference path, a meta-path instantiation query is performed on the heterogeneous graph to return all industry chain nodes and security nodes that satisfy the meta-path template. The query process is based on the influence coefficient of the event-industry chain mapping edge and the weight of the industry chain-security mapping edge. Only when both are greater than their respective preset thresholds are the corresponding security nodes included in the target stock set. The inclusion in the target stock set also includes excluding stocks that are suspended from trading on the same day, locked at the daily price limit, or have a market capitalization lower than a preset threshold, based on the stock fundamental indicators and real-time market data pre-stored in the security graph.
[0075] Furthermore, the system is also used to implement the following functions: For each stock in the target stock set, the user's interest intensity value for the stock at the current moment is queried from the user interest decay field, and the cumulative causal intensity value is obtained from the causal inference path that generated the stock. The interest intensity value and the cumulative causal intensity value are fused using a weighted summation method to obtain the recommendation score for each stock. The weight coefficient of the interest intensity value is a preset balance factor, which is used to control the relative importance of user personalized interests and event causal inference in the recommendation decision. The balance factor is dynamically adjusted according to the user's historical behavior or the current market volatility state. When the number of historical interactions of the user is lower than a preset threshold, the balance factor is reduced to enhance the contribution of causal inference; when the market volatility index exceeds a preset volatility threshold, the balance factor is increased to enhance the user's risk aversion interest selection.
[0076] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A personalized recommendation method for securities information based on knowledge graphs, characterized in that, include: Based on the recommendation association factors of securities information, a multi-source knowledge graph is constructed, which includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. In response to users' click and close actions on information, the system extracts the user's click frequency and close duration, and constructs a user interest decay field and cognitive gap detection model to analyze the user's fine-grained interest intensity on entities and unmet information needs. Using the current moment as the trigger condition, extract recent or upcoming event nodes from the event graph, extend the causal chain along the causal edge to generate at least one causal deduction path, and map the terminal events of the causal deduction path to the industry chain graph and the securities graph in sequence to obtain the set of affected target stocks. Each stock in the target stock set is matched with the user interest decay field, and a recommendation score for each stock is calculated by combining the causal strength of the causal inference path. Recommendation results are generated based on the recommendation scores. The recommendation results include the causal deduction path, the information set corresponding to each key event in the path, and at least one stock in the target stock set.
2. The personalized securities information recommendation method based on knowledge graphs according to claim 1, characterized in that, The event graph uses events as nodes and causal and temporal relationships between events as edges. The industry chain map is based on companies, industries and products as nodes, and upstream and downstream supply relationships and competitive relationships as edges. The securities graph uses stocks, concepts, and financial indicators as nodes, and affiliation and correlation relationships as edges; the user behavior graph uses users, information, click behavior, and close behavior as nodes, and click behavior edges and close behavior edges to record the user's click timestamp and dwell time.
3. The personalized securities information recommendation method based on knowledge graphs according to claim 1, characterized in that, In response to user clicks and close actions on information, the system extracts the user's click frequency and close duration to construct a user interest decay field and cognitive gap detection model, including: In response to a user's click on any piece of information, record the unique identifier of the clicked information, the click timestamp, and at least one entity associated with the corresponding information in the knowledge graph; In response to the user's action of closing this information, the closing timestamp is recorded, and the dwell time is calculated based on the difference between the closing timestamp and the click timestamp; Organize all historical interaction records of the same user into an interaction sequence in chronological order; For each click in the interaction sequence, based on the deviation between its dwell time and the preset expected dwell time, and combined with the logarithmic compression transformation of the dwell time, the instantaneous interest intensity value corresponding to this click is generated through a learnable nonlinear mapping. Based on the interaction sequence and instantaneous interest intensity value, a user interest decay field is constructed. The user interest decay field takes the user, entity and time point as independent variables and outputs the user's cumulative interest intensity on the specified entity at the current time. Furthermore, the ratio of users' click frequency to average dwell time on the same topic is statistically analyzed, and combined with the information entropy increment of new information under the corresponding topic, the cognitive gap of users on this topic is quantified, and a cognitive gap detection model is constructed.
4. The personalized securities information recommendation method based on knowledge graphs according to claim 3, characterized in that, The ratio of user click frequency to average dwell time on the same topic is statistically analyzed. Combined with the information entropy increment of new information under the corresponding topic, this quantifies the user's cognitive gap regarding the topic, and a cognitive gap detection model is constructed, including: Define the concept nodes or information clusters in the knowledge graph as topics, and count the total number of clicks and average dwell time of users for each topic within a preset time window. Calculate the information entropy increment of each topic within the current time window to reflect the average amount of new information carried by newly published information on this topic. Divide the total number of clicks by the sum of the average dwell time and the minimum constant, and multiply by the information entropy increment to obtain the cognitive gap index for the corresponding topic; When the cognitive gap index exceeds a preset threshold, it is determined that there is an unmet cognitive need for the corresponding topic, and a quantitative cognitive gap index is output as the output of the cognitive gap detection model.
5. The personalized securities information recommendation method based on knowledge graphs according to claim 1, characterized in that, Using the current moment as the trigger condition, extract recent or upcoming event nodes from the event graph, extend the causal chain along the causal edges, and generate at least one causal deduction path, including: Using the current time as the time reference point, historical event nodes with timestamps occurring before the current time within a first preset time window are selected from the event graph, and predicted event nodes whose probability of occurrence in the second preset time window exceeds a preset probability threshold are determined by the time series prediction model. The historical event nodes and predicted event nodes are used together as a seed event node set. Each directed causal edge in the event graph is pre-associated with a causal strength weight, which is determined based on historical time series data; Taking each node in the seed event node set as the starting node, a heuristic graph search algorithm is used to perform multi-hop expansion along the directed direction of the causal edge. In each hop expansion, only outgoing edges with causal strength weight greater than the preset edge weight threshold are selected for traversal, and the cumulative causal strength from the starting node to the current node is recorded. When the expansion depth reaches the preset maximum inference depth, or when the current node does not have an outgoing edge that satisfies the edge weight threshold, the expansion of this path is terminated, the current node is taken as the end node, and the sequence of nodes and directed edges traversed from the start node to the end node is recorded as a candidate causal inference path. All candidate causal inference paths are sorted from high to low according to their cumulative causal strength. At least one of the top-ranked paths is selected as the final causal inference path output. Each path is accompanied by its cumulative causal strength value and the sequence of event nodes involved in the path.
6. The personalized securities information recommendation method based on knowledge graphs according to claim 5, characterized in that, By mapping the terminal events of the causal deduction path sequentially to the industry chain map and security map, a set of affected target stocks is obtained, including: Obtain the terminal event node of at least one causal deduction path, wherein the terminal event node is a specific market event located at the end of the causal chain in the event graph; By using a pre-constructed set of event-industry chain mapping edges, the terminal event node is mapped to one or more industry nodes or product nodes in the industry chain graph that are affected by the event, thus obtaining a first intermediate mapping set. By using a pre-constructed set of industry chain-securities mapping edges, each industry node or product node in the first intermediate mapping set is mapped to one or more corresponding stock nodes in the securities graph to obtain the second intermediate mapping set. The duplicate stock nodes in the second intermediate mapping set are deduplicated and merged to obtain the affected target stock set.
7. The personalized securities information recommendation method based on knowledge graphs according to claim 6, characterized in that, The pre-built set of event-industry chain mapping edges includes: For each event type node in the event graph, the affected industrial chain links are predefined, and a directed mapping edge is established from the event node to the industry node or product node in the industrial chain graph. Each mapping edge is accompanied by an influence coefficient, which represents the transmission strength of this event to the target industry or product. The mapping edge and its influence coefficient are obtained by combining one or more of the following methods: regression analysis based on industry index fluctuations after historical events, domain expert knowledge annotation, or automatic extraction of event-industry co-occurrence patterns from research reports and news using natural language processing technology to quantify their correlation strength.
8. The personalized securities information recommendation method based on knowledge graphs according to claim 6, characterized in that, Both obtaining the first intermediate mapping set and obtaining the second intermediate mapping set employ a cross-graph linkage query mechanism, which includes: Event nodes, industry chain nodes, and securities nodes are regarded as different types of vertices in a heterogeneous graph. Cross-graphite path templates are predefined, which sequentially include event node type, industry chain node type, and securities node type. For the terminal event node of each causal deduction path, perform a meta-path instantiation query on the heterogeneous graph to return all industry chain nodes and securities nodes that satisfy the meta-path template. The query process is based on the influence coefficient of the event-industry chain mapping edge and the weight of the industry chain-securities mapping edge. Only when both are greater than their respective preset thresholds will the corresponding securities node be included in the target stock set. The target stock set also includes stocks that are suspended from trading on the same day, locked at the daily price limit, or whose market capitalization is below a preset threshold, based on the stock fundamental indicators and real-time market data pre-stored in the securities chart.
9. The personalized securities information recommendation method based on knowledge graphs according to claim 1, characterized in that, Each stock in the target stock set is matched with the user interest decay field, and a recommendation score is calculated for each stock based on the causal strength of the causal inference path, including: For each stock in the target stock set, query the user's interest intensity value for this stock at the current moment from the user interest decay field, and obtain the cumulative causal intensity value from the causal inference path that generated this stock. The interest intensity value and the cumulative causal intensity value are fused using a weighted summation method to obtain a recommendation score for each stock; The weighting coefficient of the interest intensity value is a preset balance factor, which is used to control the relative importance of user personalized interests and event causal inference in recommendation decisions. The balance factor is dynamically adjusted according to the user's historical behavior or the current market volatility. When the number of the user's historical interactions is lower than a preset threshold, the balance factor is reduced to enhance the contribution of causal inference. When the market volatility index exceeds a preset volatility threshold, the balance factor is increased to enhance the user's risk aversion interest selection.
10. A personalized securities information recommendation system based on knowledge graphs, characterized in that: The system is used to execute the knowledge graph-based personalized securities information recommendation method as described in any one of claims 1-9, and the system includes: The graph construction module is used to construct a multi-source knowledge graph based on recommendation association factors of securities information. The multi-source knowledge graph includes at least an event graph, an industry chain graph, a securities graph, and a user behavior graph. The model building module is used to respond to users' click and close operations on information, extract the user's click frequency and close duration, and build a user interest decay field and cognitive gap detection model to analyze the user's fine-grained interest intensity on entities and unmet information needs. The path generation module is used to extract recent or upcoming event nodes from the event graph, using the current time as the trigger condition, extend the causal chain along the causal edge, generate at least one causal inference path, and map the terminal events of the causal inference path to the industry chain graph and the securities graph in sequence to obtain the set of affected target stocks. The recommendation score calculation module is used to match each stock in the target stock set with the user interest decay field, and calculate the recommendation score for each stock by combining the causal strength of the causal inference path. The recommendation result generation module is used to generate recommendation results based on the recommendation score. The recommendation results include the causal inference path, the information set corresponding to each key event of the path, and at least one stock in the target stock set.