Large model content dynamic assembly method and system driven by multi-source heterogeneous field data
By performing spatiotemporal alignment and feature calibration on multi-source heterogeneous data in the financial field, and utilizing spatiotemporal graph neural networks and financial event knowledge graphs, we have solved the problems of integration efficiency and accuracy of multi-source heterogeneous data, achieved efficient and accurate data fusion and risk calibration, and provided personalized content assembly strategies.
Patent Information
- Application Number
- CN202511095102.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing multi-source heterogeneous data in the financial field, existing technologies have difficulty in efficiently and stably obtaining comprehensive data, and are unable to accurately perform spatiotemporal alignment processing, resulting in low data fusion quality and affecting the accuracy of subsequent analysis.
By acquiring structured, semi-structured and unstructured data in the financial field, performing spatiotemporal alignment processing, using spatiotemporal graph neural networks to extract multimodal features, generating multi-source feature embedding tensors, and generating feature vectors by locating key nodes, combining with the financial event knowledge graph for feature calibration, generating risk calibration feature tensors, and finally outputting personalized content assembly strategies.
It achieves precise spatiotemporal calibration across data sources, improves data integration efficiency and accuracy, can truly reflect the intrinsic correlation between data, provide a high-quality data foundation, provide accurate feature basis for subsequent decision-making, and can dynamically adjust to adapt to market risks and user needs.
Smart Images

Figure CN120597215A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for dynamically assembling large model content driven by multi-source heterogeneous field data. Background Art
[0002] Existing data processing technologies have several drawbacks when dealing with multi-source, heterogeneous data in the financial sector. During data acquisition, significant differences exist in data update frequency, data formats, and data interfaces across different data sources, making it difficult to efficiently and stably acquire comprehensive data. For example, real-time trading data in some financial markets is updated at milliseconds, while corporate financial statement data is updated quarterly or annually. This difference in update frequency makes synchronous data acquisition difficult. Furthermore, structured data is stored in relational databases, semi-structured data is often found in files with specific formats, and unstructured data is scattered across various documents and online platforms. These varying data formats significantly increase the technical complexity of data acquisition. Existing technologies are unable to accurately align multi-source data in terms of spatiotemporal alignment prior to data fusion. Financial data exhibits distinct temporal and spatial characteristics, such as varying trading hours in financial markets across different regions and the impact of economic events that are transmitted across time and space. However, some current technologies struggle to comprehensively account for these factors, resulting in low-quality fused datasets and an inability to truly reflect the inherent connections between the data. For example, when analyzing the business data of multinational financial institutions, existing spatiotemporal alignment methods are unable to accurately match relevant data due to temporal differences across different countries and regions and the spatial characteristics of business operations, impacting the accuracy of subsequent analysis. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method and system for dynamically assembling large model content driven by multi-source heterogeneous field data, which greatly improves the efficiency and accuracy of data integration.
[0004] In order to solve the above technical problems, the technical solutions of the present invention are as follows: In a first aspect, a method for dynamically assembling large model content driven by multi-source heterogeneous domain data includes: Step S1: Obtain structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment processing; Step S2: Based on the fused dataset, multimodal features are extracted through the spatiotemporal graph neural network, and multi-source feature embedding tensors are output through dynamic weight fusion; Step S3: In the high-dimensional space of the multi-source feature embedding tensor, generate feature vectors by locating key nodes. That is, capture the long-term trend vectors of macroeconomic indicators in the structured feature aggregation area, and capture the short-term mutation vectors of multi-source text sentiment signals in the unstructured feature aggregation area. Based on the feature vectors, connect the long-term trend vectors and the short-term mutation vectors to generate a feature calibration axis. Calculate the semantic projection length of the feature calibration axis in the financial event knowledge graph. Step S4: The spatial angle between the axis and the preset ideal decision plane is measured based on the semantic projection length. When the spatial angle exceeds a dynamic threshold, a numerical correction is generated based on the market volatility and the user risk factor, and the multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor. Step S5: Convert the risk calibration feature tensor into a personalized content assembly strategy based on the user profile; Step S6: Based on the personalized content assembly strategy and risk calibration feature tensor, the neural symbolic rule engine is used to perform rule-constrained content generation, and an explainable investment report containing investment portfolio plans, data traceability items, and dynamic adjustment instructions is output.
[0005] Secondly, a large-scale model content dynamic assembly system driven by multi-source heterogeneous domain data includes: The acquisition module is used to acquire structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment processing; The multi-source module is used to extract multimodal features based on the fused dataset through the spatiotemporal graph neural network, and output the multi-source feature embedding tensor through dynamic weight fusion; The calibration module is used to generate feature vectors by locating key nodes in the high-dimensional space of the multi-source feature embedding tensor. Specifically, it captures the long-term trend vectors of macroeconomic indicators in the structured feature aggregation area and the short-term mutation vectors of multi-source text sentiment signals in the unstructured feature aggregation area. Based on the feature vectors, it connects the long-term trend vectors and the short-term mutation vectors to generate a feature calibration axis. The semantic projection length of the feature calibration axis in the financial event knowledge graph is calculated. The risk calibration module is used to measure the spatial angle between the axis and the preset ideal decision plane based on the semantic projection length. When the spatial angle exceeds the dynamic threshold, a numerical correction is generated based on market volatility and user risk coefficient, and a multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor. A conversion module, which converts the risk calibration feature tensor into a personalized content assembly strategy based on the user profile; The engine module is used to assemble strategies and risk-calibrated feature tensors based on personalized content, perform rule-constrained content generation through a neural symbolic rule engine, and output an explainable investment report containing investment portfolio plans, data traceability entries, and dynamic adjustment instructions.
[0006] According to a third aspect, a computing device includes: one or more processors; The storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method.
[0007] In a fourth aspect, a computer-readable storage medium stores a program, which implements the method when executed by a processor.
[0008] The above solution of the present invention includes at least the following beneficial effects: In view of the different characteristics of structured, semi-structured and unstructured data, accurate spatiotemporal calibration across data sources is achieved, ensuring that the generated fused data set can truly reflect the intrinsic correlation between the data, laying a high-quality data foundation for subsequent data processing, and greatly improving the efficiency and accuracy of data integration.
[0009] At the feature extraction level, multimodal features are extracted using a spatiotemporal graph neural network, and combined with dynamic weight fusion to output a multi-source feature embedding tensor, significantly improving the comprehensiveness and accuracy of feature extraction. This approach fully captures the unique characteristics of different data types and their complex relationships. The dynamic weighting setting adaptively adjusts the importance of each source feature based on the data's characteristics, making the extracted features more reflective of the essential laws of financial data and overcoming the shortcomings of traditional methods in insufficiently mining multimodal features.
[0010] In terms of feature analysis and calibration, key nodes are located to capture long-term trend vectors and short-term mutation vectors, generate feature calibration axes, and calculate their semantic projection lengths, enabling precise analysis of key information in high-dimensional feature spaces. This approach simultaneously considers long-term trends in macroeconomic indicators and short-term variations in sentiment signals from multiple text sources. Combined with the semantic projection of the financial event knowledge graph, this approach enhances feature analysis depth and semantic relevance, providing a more accurate feature basis for subsequent decision-making.
[0011] During the risk calibration phase, based on spatial angle determination and a dynamic threshold mechanism, market volatility and user risk factors are combined to generate corrections, which are then injected into the feature tensor, achieving dynamic risk calibration of feature data. This process adjusts feature data in real time based on market changes and user risk preferences, making the risk calibration feature tensor more aligned with actual market risk conditions and user needs, and enhancing the adaptability of data processing to dynamic risk environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of a method for dynamically assembling large model content driven by multi-source heterogeneous domain data provided by an embodiment of the present invention.
[0013] Figure 2 This is a schematic diagram of a large model content dynamic assembly system driven by multi-source heterogeneous domain data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0015] like Figure 1 As shown, an embodiment of the present invention proposes a method for dynamically assembling large model content driven by multi-source heterogeneous domain data, the method comprising the following steps: Step S1: Obtain structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment processing; Step S2: Based on the fused dataset, multimodal features are extracted through the spatiotemporal graph neural network, and multi-source feature embedding tensors are output through dynamic weight fusion; Step S3: In the high-dimensional space of the multi-source feature embedding tensor, generate feature vectors by locating key nodes. That is, capture the long-term trend vectors of macroeconomic indicators in the structured feature aggregation area, and capture the short-term mutation vectors of multi-source text sentiment signals in the unstructured feature aggregation area. Based on the feature vectors, connect the long-term trend vectors and the short-term mutation vectors to generate a feature calibration axis. Calculate the semantic projection length of the feature calibration axis in the financial event knowledge graph. Step S4: The spatial angle between the axis and the preset ideal decision plane is measured based on the semantic projection length. When the spatial angle exceeds a dynamic threshold, a numerical correction is generated based on the market volatility and the user risk factor, and the multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor. Step S5: Convert the risk calibration feature tensor into a personalized content assembly strategy based on the user profile; Step S6: Based on the personalized content assembly strategy and risk calibration feature tensor, the neural symbolic rule engine is used to perform rule-constrained content generation, and an explainable investment report containing investment portfolio plans, data traceability items, and dynamic adjustment instructions is output.
[0016] In an embodiment of the present invention, by fusing structured, semi-structured, and unstructured data and achieving spatiotemporal alignment, the problem of single data is effectively eliminated, providing solid data support for subsequent feature extraction and decision-making. By extracting multimodal features and dynamically weighting them with the help of a spatiotemporal graph neural network, the system can accurately capture the long-term trends of macroeconomic indicators and discern short-term changes in sentiment across multiple sources. Furthermore, by semantically projecting and associating financial event knowledge graphs, the system enhances the explanatory power of features for market dynamics. By measuring the angle between the feature axis and the ideal decision plane in real time and combining market volatility and user risk factors to generate corrections, the feature tensor can be dynamically adjusted to adapt to changes in risk. When the market fluctuates or user risk preferences change, corrections can be injected in a timely manner to make the risk calibration feature tensor more aligned with the actual risk scenario. By converting the risk calibration feature tensor into a personalized content assembly strategy based on user profiles, the final output content can be precisely matched to the risk preferences, investment objectives, and information needs of different users. A neural symbolic rule engine generates a report containing investment portfolio plans, data traceability, and dynamic adjustment instructions, making the decision-making process transparent and traceable through data traceability and adjustment instructions.
[0017] In a preferred embodiment of the present invention, the above step S1: obtaining structured data, semi-structured data and unstructured data in the financial field, and generating a fused data set through spatiotemporal alignment processing, includes: Step S11, acquiring structured, semi-structured and unstructured original data in the financial field to obtain an original heterogeneous data set; Step S12: Based on the original heterogeneous data set, missing value filling and outlier smoothing are performed on the structured data, table topology parsing and semantic conflict elimination are performed on the semi-structured data, and noise filtering and sentiment calibration are performed on the structured data to obtain a preliminary screening quality data set; Step S13: Based on the initial screening quality dataset, the financial event knowledge graph is used to drive spatiotemporal alignment, that is, the timestamps of low-frequency economic indicators and high-frequency event data streams are aligned in the time dimension; in the spatial dimension, the enterprise addresses and geographic tag data are mapped to a unified coordinate grid to obtain a spatiotemporal correlation dataset; In step S14, a modal unified conversion is performed on the spatiotemporal correlation dataset, that is, the numerical data is normalized to the interval [0, 1], the text data is encoded into a semantic vector of a preset dimension, and the tabular data is converted into a graph node embedding vector to generate a fused dataset.
[0018] In an embodiment of the present invention, by synchronously acquiring structured, semi-structured and unstructured data, an original heterogeneous data set is constructed to make the acquired data more comprehensive; missing value filling and outlier smoothing for structured data can reduce the interference of numerical fluctuations, topological analysis and semantic conflict elimination of semi-structured data can solve format confusion and expression contradictions, and noise filtering and sentiment calibration of unstructured data can improve the effectiveness of text information; multiple processing greatly improves the efficiency of data integration; relying on the financial event knowledge graph to achieve time and space dual-dimensional alignment, unify the timestamps of low-frequency economic indicators and high-frequency event data in time, and avoid the association break caused by time granularity differences; unify the coordinate system of enterprise addresses and geographic markers in space, and eliminate regional information confusion; through modal unified conversion, heterogeneous data such as numbers, texts, and tables are converted into standardized formats, eliminating format barriers of different data types.
[0019] In the embodiments of the present invention, when applied specifically, it can be achieved through the following technical solutions, for example: In step S11 above, three types of data are collected in real time from financial data platforms (such as Wind, Eastmoney, and Reuters): Structured data: including macroeconomic indicators (GDP, CPI, etc., updated daily / monthly), financial market data (stock prices, trading volume, etc., updated minute-by-minute), and corporate financial data (revenue, profit, etc., updated quarterly / annually); Semi-structured data: This includes table fragments in corporate annual reports (such as consolidated balance sheets), nested lists in industry reports (such as market share rankings), and form data in regulatory documents (such as responses to inquiry letters from listed companies). Unstructured data: includes news information (such as policy interpretation articles), social media comments (such as stock forum discussions), and full text of corporate announcements (such as major contract disclosures).
[0020] The collected raw data is classified and stored according to "data type-source-timestamp", forming an original heterogeneous data set containing millions of entries (such as 100,000 structured indicators, 500,000 semi-structured tables, and 5 million unstructured texts), retaining the original data format (such as CSV, PDF, TXT) and metadata (such as collection time, data credibility score).
[0021] In the above step S12, structured data processing: missing values in the time series (such as no data on a certain trading day due to holidays) are filled with the mean of the three periods before and after; non-time series data (such as the number of employees in a company) are filled with the mean of the same industry (for example, when the number of employees of a certain technology company is missing, it is replaced by the mean of five comparable companies in the industry); abnormal fluctuation points are identified by the 3σ rule (such as a single-day stock price fluctuation of more than 10% and deviating from the moving average by 3 times the standard deviation), and local weighted regression (LOWESS) is used for smoothing correction to retain the trend while eliminating extreme value interference (such as replacing abnormal points with the weighted average of the five periods before and after).
[0022] Semi-structured data processing: For nested tables (such as the "division report" in the annual report containing sub-tables), a topological structure is constructed by identifying the hierarchical relationship of the table headers (such as "parent company data" and "merged data"), and the associated fields of the parent and child tables are marked (such as the proportion of "division income" to "total revenue"); for different names of the same indicator (such as "operating income" and "main business income"), the field names are unified through the financial field synonym library; for unit conflicts (such as "10,000 yuan" and "100 million yuan"), batch conversion is carried out according to 100 million yuan = 100 million yuan.
[0023] Unstructured data processing: Identify and delete advertising text through keyword matching (such as "advertisement," "promotion," and "scan the code to follow"); retain the earliest published version for duplicate content (such as multiple reposts of the same news); re-score text with shifted sentiment (such as "The company's performance has been steadily increasing, but the growth rate is lower than expected") based on financial sentiment dictionaries (such as the Loughran-McDonald dictionary), and correct deviations caused by general sentiment lexicons (such as "liability" is a neutral word in the financial context, not a negative one).
[0024] In step S13, time dimension alignment is performed to address the time stamp differences between low-frequency data (e.g., monthly CPI) and high-frequency text event stream data. The window size is adaptively adjusted based on the data frequency: low-frequency indicators (monthly) are aligned to a 30-day window, and daily text events are aligned by month. Medium-frequency indicators (daily) are aligned to a 1-day window, and minute-level stock price data is aggregated into daily averages. For fuzzy timestamp data (e.g., "this week's policy"), contextual analysis is used to map it to precise timestamps. Spatial dimension alignment: The registered address of the enterprise is converted into longitude and latitude coordinates through geocoding, and the geographic label of the text event (such as "Yangtze River Delta Region") is mapped to the central coordinate grid. Spatial association is established by grouping by coordinate grid: Enterprise entities within the same grid are bound to regional text events. For example, new energy enterprises are spatially associated with the "Yangtze River Delta New Energy Policy" text event.
[0025] In the above step S14, the normalization of numerical data is to compress the numerical data such as macroeconomic indicators and stock prices by using Min-Max normalization to compress the data of different dimensions (such as "%" of CPI and "yuan" of stock prices) to the range of [0, 1]. For example, the GDP growth rate (-5%~15%) is mapped to a value of 0~1 (-5% corresponds to 0, 15% corresponds to 1).
[0026] Text data encoding uses a pre-trained word embedding model in the financial field (such as a financial fine-tuning model based on BERT) to convert text data into a 300-dimensional semantic vector. For example, for the sentence "The central bank lowered the reserve requirement ratio by 0.5 percentage points", the encoded vector needs to highlight the semantic weight of financial keywords such as "lower the reserve requirement ratio" and "central bank".
[0027] Tabular data conversion is to convert semi-structured tables into graph node embedding vectors. For example, the "supplier list" in a company's annual report can be converted into an "company-supplier" relationship graph. Each company is a node, and the node embedding vector contains attributes such as transaction amount and years of cooperation. A 300-dimensional embedding vector is generated through a graph neural network.
[0028] The converted numerical vectors, text semantic vectors, and graph node vectors are spliced according to the "time-space" dimension to form a fused dataset in a unified format. Each data piece contains "space-time identifier + multimodal feature vector" (such as "2023-07-01_Shanghai_31.23°N_121.50°E+[0.6 (GDP vector),...,0.3 (text semantic feature)]").
[0029] In a preferred embodiment of the present invention, the above step S2: extracting multimodal features based on the fused dataset through a spatiotemporal graph neural network, and outputting a multi-source feature embedding tensor through dynamic weight fusion, includes: Step S21: Input the structured data in the fusion dataset into the time series convolution layer to capture the cyclical evolution of macroeconomic indicators and output the numerical evolution feature vector; Step S22: Input the semi-structured data in the fused dataset into the graph attention network to analyze the equity linkage chain and supply chain topology of the corporate financial report and industry report, and output the inter-table linkage feature vector; Step S23: input the unstructured data in the fusion dataset into a multi-head text encoder, calculate the text sentiment extreme value based on the weighted financial sentiment dictionary, and output the text semantic feature vector; In step S24, the numerical evolution feature vector, the inter-table association feature vector, and the text semantic feature vector are input into the dynamic weight allocation unit. By real-time scanning of the policy nodes and market fluctuation signals in the financial event knowledge graph, the preset weight allocation rules are matched according to the event type, and the numerical evolution feature vector, the inter-table association feature vector, and the text semantic feature vector are weightedly spliced and tensor regularized to output a multi-source feature embedding tensor.
[0030] In an embodiment of the present invention, structured data captures the periodic evolution law through the time series convolution layer, avoiding the omission of periodic signals by traditional time series models; semi-structured data uses the graph attention network to analyze equity relationships and supply chain topology, which can explore hidden business linkages between enterprises and break through the limitations of surface information of tabular data; unstructured data is combined with the financial sentiment dictionary through a multi-head text encoder to avoid the misjudgment of financial professional semantics by general text models; through the combination of numerical evolution features, inter-table association features and text semantic features, the three are combined to cover the "macro-meso-micro" full dimensions of the financial market; dynamic weight allocation avoids "one-size-fits-all" feature fusion, making the collaboration of multi-source features more in line with real-time market logic; after dynamic weight weighting and tensor regularization, the output multi-source feature embedding tensor has strong correlation, high adaptability and excellent availability.
[0031] In the embodiments of the present invention, when applied specifically, it can be achieved through the following technical solutions, for example: In the above step S21, the structured data in the fusion data set (such as time series data of macroeconomic indicators such as GDP growth rate, CPI, interest rate, etc., updated on a daily / monthly / quarterly basis) are integrated.
[0032] Multi-scale convolution kernels (such as 3, 6, and 12 time steps) are used to perform sliding convolution on time series data: small-scale convolution kernels (3 steps) capture short-term fluctuations (such as the month-on-month changes in monthly CPI); large-scale convolution kernels (12 steps) capture long-term cycles (such as the annual growth trend of GDP).
[0033] Residual connections are used to alleviate gradient vanishing and retain the characteristics of key time nodes (such as the jump in economic indicators in the month when the policy is released); a numerical evolution feature vector is generated, which contains information such as the short-term fluctuation amplitude, long-term trend slope, cycle peak position, etc. of macroeconomic indicators (such as "GDP has grown for five consecutive quarters with a slope of 0.02; CPI quarterly fluctuation standard deviation is 0.3").
[0034] The above step S22 integrates the semi-structured data in the dataset (such as the "equity structure table" and "supply chain cooperation list" in the company's annual report, and the "company competition landscape table" in the industry report).
[0035] Enterprises and industries are regarded as nodes, equity associations (such as "holding more than 51% is a controlling stake") and supply chain relationships (such as "core suppliers" and "major customers") are regarded as edges, and the edge weight is the association strength (such as shareholding ratio and transaction amount ratio). For each node (such as "Enterprise A"), the attention weight of it and the adjacent nodes (such as "shareholder B" and "supplier C") is calculated: the higher the association strength (such as Enterprise A holds 60% of the equity of Enterprise D), the greater the attention weight (such as 0.8); for non-core associations (such as the transaction proportion between Enterprise A and Enterprise E is only 5%), the attention weight is smaller (such as 0.2).
[0036] The node's own characteristics (such as the revenue scale of enterprise A) and the characteristics of its high-weight neighbors (such as the industry influence of shareholder B) are weighted and fused to generate a node embedding vector; the inter-table association feature vector is generated, which contains topological information such as the enterprise's equity control chain, supply chain dependence, and industry competition landscape (such as "Enterprise A's core supplier is C, with a dependence of 0.7; it is controlled by shareholder B, with a control of 0.6").
[0037] In step S23 above, the unstructured data in the dataset (such as policy documents, news reports, corporate announcements, and other text data) is integrated; the text is segmented according to financial terms (such as "RRR cut", "merger and acquisition", and "default"), and converted into basic semantic vectors through the pre-trained financial BERT model.
[0038] Set 3-6 attention heads to focus on different dimensions of the text: The first one focuses on policy keywords (such as "central bank" and "monetary policy"); The first two focus on market impact descriptions (e.g., “stock prices increased,” “volatility increased”); The first three focus on sentiment-oriented words (such as "positive", "risk", and "warning").
[0039] Based on the financial sentiment dictionary (such as those containing labels such as "good news" and "bad news"), the attention weight is weighted to calculate the overall sentiment extreme value of the text (such as "the sentiment score of the policy document is +0.8, which is strongly positive; the sentiment score of the default news is -0.7, which is strongly negative"); a text semantic feature vector is generated, which contains information such as the core semantics of the text, sentiment polarity, and keyword weight (such as "the text focuses on the 'reserve requirement ratio reduction policy', with positive sentiment accounting for 0.8 and the associated 'market liquidity' weight being 0.6").
[0040] The above step S24: based on the numerical evolution feature vector, the inter-table association feature vector, and the text semantic feature vector, the policy nodes (such as "central bank cuts the reserve requirement ratio" and "interest rate hikes") and market volatility signals (such as "stock market volatility ≥ 5%" and "bond market yield jumps") in the knowledge graph are retrieved in real time (such as every 5 minutes).
[0041] If a "policy release" event (such as a reserve requirement ratio cut) is detected, the preset rules are matched: the weight of the numerical evolution feature vector is increased (such as 0.6), the weight of the text semantic vector is reduced (such as 0.2), and the weight of the inter-table association vector is moderate (such as 0.2); If a "severe market fluctuation" event (such as a stock market crash) is detected, the matching rule is: increase the weight of the text semantic vector (such as 0.5), reduce the weight of the numerical evolution vector (such as 0.3), and the weight of the inter-table association vector is 0.2; Under normal market conditions, balanced weights are used (0.4 for numerical values, 0.3 for table values, and 0.3 for text values).
[0042] According to the above weights, a weighted sum is performed on the numerical evolution feature vector, the inter-table association feature vector, and the text semantic feature vector (for example, in the case of policy events: 0.6×numerical vector + 0.2×inter-table vector + 0.2×text vector); L2 regularization is performed on the concatenated vector (the vector modulus is scaled to 1) to prevent a certain type of feature from dominating the tensor due to excessively large numerical values, and a multi-source feature embedding tensor with a unified dimension (such as 512 dimensions) is generated.
[0043] The spatiotemporal graph neural network is the core architecture for feature extraction. It integrates the spatiotemporal characteristics of multimodal data through a three-layer structure of "temporal component + spatial component + fusion layer": The time component is a time series convolution layer, which processes the time series characteristics of structured data and captures the cyclical evolution of macroeconomic indicators (such as seasonality and cyclical fluctuations). It consists of three layers of convolution blocks, each of which contains a convolution layer (multi-scale convolution kernel), a batch normalization layer, and a ReLU activation function, and transmits long-term dependencies (such as cross-year economic trends) through residual connections.
[0044] The spatial component is a graph attention network, which processes the spatial correlation characteristics of semi-structured data and analyzes the topological relationships between enterprises and industries (such as equity control and supply chain dependence). It consists of two layers of graph attention, each of which performs weighted aggregation on the features of node neighbors (attention weights are dynamically calculated based on the strength of the association), highlighting core associations (such as the influence of the parent company on its subsidiaries) and weakening marginal associations (such as minority shareholders' holdings).
[0045] The fusion layer integrates multimodal features, fusing the temporal features of the time component, the topological features of the spatial component, and the text semantic features through a cross-attention mechanism to capture the cross-modal association of "policy-enterprise-market" (such as reserve requirement ratio reduction policy → increased bank liquidity → reduced financing costs for supply chain enterprises); the three types of features are mapped to the same dimensional space through the fully connected layer, the mutual information between features is calculated (such as the correlation between numerical features and text features), the fusion weights are dynamically adjusted, and finally a high-dimensional (such as 1024-dimensional) multi-source feature embedding tensor is output.
[0046] In a preferred embodiment of the present invention, the above step S3: in the high-dimensional space of the multi-source feature embedding tensor, generating a feature vector by locating key nodes, that is, capturing the long-term trend vector of the macroeconomic indicator in the structured feature aggregation area, and capturing the short-term mutation vector of the multi-source text sentiment signal in the unstructured feature aggregation area; based on the feature vector, connecting the long-term trend vector and the short-term mutation vector to generate a feature calibration axis; calculating the semantic projection length of the feature calibration axis in the financial event knowledge graph, including: Step S31: Based on the high-dimensional spatial geometric structure of the multi-source feature embedding tensor, the long-term trend vector of the structured data is captured in the first feature area and the macroeconomic trend node is located. At the same time, the short-term mutation vector of the unstructured data is captured in the second feature area and the text sentiment signal mutation node is located. Step S32, connecting the macroeconomic trend node and the text sentiment signal mutation node to form a feature calibration axis; Step S33, calculating the direction vector modulus of the characteristic calibration axis as the axis reference length, and decomposing the axis direction vector into policy dimension and market dimension components to generate a two-dimensional direction identifier; Step S34, using the two-dimensional direction identifier as the search key, retrieve the event node sequence matching the policy market label in the financial event knowledge graph, calculate the projection scalar of the feature calibration axis vector in the conduction chain direction of the event node sequence, and take the accumulated value of the shortest path weight between the conduction chain nodes as the semantic projection length.
[0047] In an embodiment of the present invention, a first feature area and a second feature area are distinguished in a high-dimensional space to capture long-term trend vectors and short-term mutation vectors, respectively. This separation process avoids the long-term trend from being masked by short-term noise, and also prevents the short-term mutation from being diluted by long-term inertia, so that the key signals of two different time scales can be clearly presented. By connecting the macroeconomic trend node and the text sentiment mutation node to form a feature calibration axis, the originally independent long-term trend and short-term mutation are transformed into an organically related "trend-mutation" linkage relationship. This relationship breaks through the limitations of single feature analysis and is closer to the actual operation logic. The feature calibration axis is decomposed into policy dimension and market dimension components, and the core source of the influence can be clearly distinguished through the two-dimensional direction identification. The financial event knowledge graph is retrieved through the two-dimensional direction identification, the abstract feature calibration axis is associated with the specific event node sequence, and the semantic projection length is calculated. The vector features in the high-dimensional space can be bound to actual financial events and policy transmission paths, which greatly improves the interpretability of the feature vector to market dynamics.
[0048] In the embodiments of the present invention, when applied specifically, it can be achieved through the following technical solutions, for example: In the above-mentioned step S31, high-dimensional vectors of structured features (such as macroeconomic indicators such as GDP and CPI) in the multi-source feature embedding tensor are density clustered (such as the DBSCAN algorithm) to identify the aggregation area with dense eigenvalues (the first feature area), which represents the typical distribution of macroeconomic indicators; within the aggregation area, time series smoothing (such as a 12-period moving average) is performed on the numerical evolution feature vector output by the time series convolution layer to filter short-term fluctuations (such as occasional jumps in the monthly CPI), retain long-term trends (such as the annual CPI growth curve), and locate the endpoints of the smoothed vector as macroeconomic trend nodes. The direction of the vector reflects the direction of trend evolution (such as growth or recession).
[0049] Cluster high-dimensional vectors of unstructured features (such as text sentiment signals) to identify clustered areas with dense sentiment features (second feature areas). Within the clustered areas, calculate the first-order difference of the time series of text semantic feature vectors (such as the difference in sentiment polarity between two adjacent hours), and identify mutation points (such as the sentiment polarity suddenly changes from +0.6 to -0.4 after a policy is released) through the threshold method (such as the absolute value of the difference > 0.5). The endpoints of the feature vectors corresponding to the mutation points are located as the text sentiment signal mutation nodes, and the vector direction reflects the sentiment change before and after the mutation (such as from positive to negative).
[0050] In the above step S32, in the high-dimensional feature space, a straight line is used to connect the macroeconomic trend node and the text sentiment signal mutation node to form a feature calibration axis; the starting point of the axis is the macroeconomic trend node (representing the stable trend of long-term, structured features), and the end point is the text sentiment signal mutation node (representing the sudden change of short-term, unstructured features); the direction of the axis reflects the relationship between "how long-term trends are affected by short-term mutations" (such as the macroeconomic growth trend changes due to sudden negative event signals), and the length reflects the spatial distance between the two (the farther the distance, the greater the deviation between the long-term trend and the short-term mutation).
[0051] In the above step S33, the direction vector of the characteristic calibration axis is extracted (obtained by subtracting the starting point coordinates from the end point coordinates), and the modulus of the vector (i.e., the straight-line distance between two points in space) is calculated as the axis reference length; for example, if the coordinate difference between the two points in 3D space is (3, 4, 0), then the modulus is 5, which means that the axis reference length is 5.
[0052] Preset the policy dimension (corresponding to the characteristic axis related to macroeconomic policies, such as central bank interest rate adjustments and fiscal subsidies) and the market dimension (corresponding to the characteristic axis related to market performance, such as stock price volatility and trading volume); project the axis direction vector onto the policy dimension axis and the market dimension axis to obtain two projection components (such as the policy component accounts for 0.6, and the market component accounts for 0.4); use the component ratio as the dual-dimensional direction identifier (such as "policy accounts for 0.6, market accounts for 0.4"), reflecting the degree of dominance of policy factors and market factors in the axis direction (the higher the ratio, the greater the impact of the dimension on the characteristic correlation).
[0053] In the above step S34, the dual-dimensional direction identifier is used as the search key (such as "policy share ≥ 0.5 and market share ≥ 0.3") to filter event nodes in the knowledge graph that contain both policy labels (such as "RRR cut", "new regulatory rules") and market labels (such as "stock price fluctuation", "trading volume surge") (such as the event "central bank cuts RRR → stock market trading volume increases").
[0054] Through graph traversal algorithms (such as breadth-first search), the shortest path from the policy node to the market node is found (such as "the release of the reserve requirement ratio reduction policy → increased bank liquidity → capital inflow into the stock market → stock price increase"), forming a transmission chain node sequence containing 3-5 nodes.
[0055] Calculate the inner product of the characteristic calibration axis vector and the main direction vector of the conduction chain node sequence (the overall direction from the policy node to the market node) to obtain the projection scalar (the closer the value is to 1, the more consistent the axis direction is with the conduction chain logic).
[0056] The edges between each node in the transmission chain have preset weights (for example, the impact intensity of "RRR cut → increased liquidity" is 0.8, and the intensity of "increased liquidity → stock price increase" is 0.6). The weights are added up (0.8+0.6=1.4) as the semantic projection length, reflecting the intensity and complexity of policy transmission to the market.
[0057] In a preferred embodiment of the present invention, step S4: measuring the spatial angle between the axis and the preset ideal decision plane based on the semantic projection length; when the spatial angle exceeds a dynamic threshold, generating a numerical correction value based on the market volatility and the user risk factor, and injecting the multi-source feature embedding tensor along the feature calibration axis to generate a risk calibration feature tensor, includes: Step S41: generating a dynamic deviation threshold based on the semantic projection length and the real-time market volatility index; measuring the spatial angle between the feature calibration axis and the preset rational decision plane; Step S42: When the spatial angle exceeds the dynamic deviation threshold, the angle difference by which the spatial angle exceeds the threshold is calculated, the fluctuation tolerance coefficient in the user risk profile is taken, and the angle difference is multiplied by the fluctuation tolerance coefficient and then by the inverse of the semantic projection length to obtain a scalar correction value; Step S43, converting the scalar correction amount into a unit vector in the axis direction, linearly enhancing or attenuating the characteristic component in the eigenvalue tensor in the axis direction, and obtaining a corrected eigenvalue tensor; In step S44, the variance of each dimension of the feature tensor is calculated, the standard deviation of the dimension exceeding the mean of the overall variance is scaled, and the range of the scaled feature tensor value is constrained to be between zero and one, thereby generating a risk calibration feature tensor.
[0058] In an embodiment of the present invention, a dynamic deviation threshold is generated based on the semantic projection length and real-time market volatility. When the market fluctuates violently, the threshold is relaxed as volatility increases; when the market is stable, the threshold is tightened. This dynamic adjustment allows risk assessments to better align with real-time market conditions, avoiding misjudgments or missed assessments caused by static thresholds. The angle difference is associated with the "volatility tolerance coefficient" in the user's risk profile to generate a scalar correction. Users with high risk tolerance receive a smaller correction, while risk-averse users receive a larger correction. Feature components are enhanced or attenuated along the feature calibration axis. This directional correction reduces redundant interference with valid features, allowing risk adjustment to focus more on core issues and improving correction efficiency. Variance analysis is used to identify high-volatility feature dimensions, perform standard deviation scaling, and constrain the value range.
[0059] In the embodiments of the present invention, when applied specifically, it can be achieved through the following technical solutions, for example: In step S41 above, the threshold is adjusted by combining the semantic projection length (the accumulated value of the transmission chain node weights, reflecting the complexity of policy-market transmission) and the real-time market volatility index (such as the CSI 300 Volatility Index, which measures the severity of market fluctuations): If the semantic projection length is long (the transmission chain is complex, such as policy → industry → enterprise → market, which requires multiple links), the threshold can be appropriately relaxed (for example, allowing an angle of ≤15°); If the real-time volatility is high (e.g., daily volatility > 8%), the threshold is also relaxed (e.g., allowing an angle ≤ 20°) to avoid overcorrection; If the transmission is simple and the volatility is low, the threshold is tightened (for example, the angle is allowed to be ≤5°) and the deviation is strictly controlled.
[0060] Example: The semantic projection length is 5 (medium complexity), the volatility index is 3% (low volatility), and the dynamic deviation threshold is set to 8°.
[0061] It is constructed based on the feature distribution of historical optimal decision-making cases, representing the ideal state of "the highest match between long-term trends and short-term mutations"; the spatial angle between the feature calibration axis and the plane is measured: the smaller the angle, the closer the current feature association is to the ideal decision-making state; the larger the angle, the greater the degree of deviation (for example, if the angle between the axis and the plane is 12°, which is greater than the dynamic threshold of 8°, it is determined to require correction).
[0062] In the above step S42, when the spatial angle (eg, 12°) exceeds the dynamic deviation threshold (eg, 8°), the difference is calculated as 12°-8°=4°.
[0063] The volatility tolerance coefficient is extracted from the user's risk profile and reflects the user's acceptance of market fluctuations: The conservative user coefficient is 0.5 (sensitive to deviations and small correction range); The coefficient of aggressive users is 1.0 (high tolerance for deviations and large correction range).
[0064] The reciprocal of the semantic projection length is that the longer the length (the more complex the conduction chain), the smaller the reciprocal (e.g., the reciprocal of length 5 is 0.2), which reduces the correction amplitude (avoiding over-adjustment in complex conduction); Calculation example: Angle difference 4° × fluctuation tolerance coefficient 0.5 (conservative) × 0.2 (the reciprocal of length 5) = 0.4, that is, the scalar correction amount is 0.4.
[0065] In the above step S43, the scalar correction value (such as 0.4) is converted into a unit vector in the direction of the characteristic calibration axis (keeping the direction unchanged and the length normalized to 1), ensuring that the correction direction is consistent with the "long-term trend-short-term mutation" correlation direction.
[0066] Identify the characteristic components along the axis direction in the feature tensor (such as the correlation characteristics between macroeconomic trends and sudden changes in text sentiment); if the spatial angle is biased towards the policy dimension (the axis is closer to the policy characteristic axis), enhance the policy-related components (such as the weight of macroeconomic indicators) and attenuate the market-related components (such as short-term sentiment signals); when the correction amount is 0.4, the policy component is enhanced by 40% and the market component is attenuated by 40% (linear adjustment ensures the coordination of feature proportions).
[0067] In the above step S44, the variance of each dimension of the corrected feature tensor is calculated (e.g., the variance of the policy feature dimension is 0.3, and the variance of the market feature dimension is 0.1), the mean of the overall variance is obtained (e.g., 0.2), and the dimensions whose variance exceeds the mean (e.g., the policy feature dimension 0.3>0.2) are marked as high-volatility dimensions.
[0068] Standardize highly volatile dimensions and compress their numerical range (e.g., scale the original value distribution of policy characteristics from [0, 100] to a more concentrated distribution to reduce the impact of extreme values).
[0069] Through Min-Max normalization, the values of all components are compressed to the range of [0, 1] to ensure that different dimensional characteristics (such as policy characteristics and market characteristics) participate in subsequent calculations at the same scale (such as 0 represents the minimum impact and 1 represents the maximum impact).
[0070] In a preferred embodiment of the present invention, the above step S5: converting the risk calibration feature tensor into a personalized content assembly strategy based on the user profile includes: Step S51: Orthogonally project the risk calibration feature tensor into a macroeconomic trend subspace, an industry association subspace, and a text signal pulse subspace. Simultaneously, a subspace operation constraint group is generated based on the user profile, i.e., the user risk level is mapped to a feature selection threshold in the trend subspace; the investment target is converted into a node filtering condition in the industry subspace; and the capital scale is quantified as a sensitivity attenuation coefficient in the text signal subspace. Step S52: In response to real-time market event signals, the financial regulatory rule base is activated using the subspace operation constraint group as a modulation parameter, and three types of rules are modulated: a risk matching rule acting on the macroeconomic trend subspace is modulated using a feature selection threshold, a target adaptation rule acting on the industry association subspace is expanded through a node filtering condition, and a scale response rule acting on the text signal impulse subspace is scaled using a sensitivity attenuation coefficient; Step S53: execute the modulated risk matching rule in the trend subspace to extract trend segments; execute the expanded target adaptation rule in the industry association subspace to retrieve the industry chain path; execute the scaled scale response rule in the text signal subspace to filter text event tags; and fuse the trend segments, industry chain paths and text event tags to generate a personalized content assembly strategy.
[0071] In an embodiment of the present invention, the risk calibration feature tensor is orthogonally projected into three subspaces: macroeconomic trends, industrial associations, and text signal pulses, to avoid cross-interference of features of different dimensions and improve the targeted utilization of features; the user portrait is concretized as a subspace operation constraint group, the risk level corresponds to the feature selection threshold of the trend subspace, the investment target is converted into the node filtering condition of the industry subspace, and the capital scale is quantified as the sensitivity attenuation coefficient of the text signal. This conversion allows abstract user needs to become executable quantitative parameters. In response to real-time market event signals, three types of rules are modulated with constraint groups as parameters: risk matching rules are dynamically adjusted according to user risk thresholds, target adaptation rules expand the industry chain screening range according to investment goals, and scale response rules scale text signal sensitivity according to fund size; dynamic modulation allows rules to adapt to real-time market changes and meet regulatory compliance; after the modulated rules are executed in each subspace, the extracted trend fragments, industry chain paths, and text event labels are all screened by personalized constraints and real-time rules, and have the dual attributes of adapting to user needs and adapting to market dynamics; the personalized content assembly strategy generated by the fusion of the three covers the stability of long-term trends, the structural nature of industry linkages, and the flexibility of short-term signals, greatly improving the adaptability of the strategy to users' actual scenarios.
[0072] In the embodiments of the present invention, when applied specifically, it can be achieved through the following technical solutions, for example: In step S51, an orthogonal projection (such as PCA orthogonal transformation) is performed on the risk calibration feature tensor to decompose the high-dimensional features into three non-interfering subspaces: Macroeconomic trend subspace: retains long-term trend characteristics related to macroeconomic indicators such as GDP and interest rates (such as the 5-year GDP growth rate curve and annual CPI change trend), and filters out short-term fluctuations; Industry-related subspace: Extracts related features such as corporate equity and supply chain (e.g., the supply chain topology of "new energy vehicle companies-battery suppliers-lithium mining companies"), focusing on node relationships within the industry; Text signal pulse subspace: retains the short-term mutation characteristics of unstructured data such as policy texts and market comments (for example, the change in text sentiment within 24 hours after a policy is released).
[0073] Subspace operation constraint group generation: Risk level → feature selection threshold: Map the user risk level (such as conservative, balanced, aggressive) to a threshold between 0 and 1: the threshold for conservative users is set to 0.8 (only macro features with a trend stability score ≥ 0.8 are retained, such as excluding industry indices with large fluctuations); the threshold for aggressive users is set to 0.5 (allowing more trend features with high volatility but great potential to be retained).
[0074] Investment target → Node filtering conditions: Convert investment targets (such as "high-dividend blue-chip stocks" and "new energy industry chain") into node attribute conditions in the industry subspace. When the target is "new energy industry chain", the condition is set to "industry classification = new energy and enterprise nature = manufacturing" to filter out non-related enterprise nodes (such as excluding financial enterprises).
[0075] Capital size → Sensitivity attenuation coefficient: A coefficient of 0.3-1.0 is generated according to the capital size (e.g., less than 100,000, more than 1 million): When the capital size is 10 million, the coefficient is set to 0.4 (reducing sensitivity to short-term text signals and avoiding frequent position adjustments of large amounts of capital due to market noise); when the capital size is 100,000, the coefficient is set to 0.9 (increasing sensitivity and facilitating small amounts of capital to quickly respond to market event signals).
[0076] In the above step S52, real-time market events (such as "central bank interest rate hike" and "new energy subsidy policy adjustment") are received through the API, and event type labels (such as "monetary policy" and "industrial policy") are generated; based on the labels, corresponding rule sets (such as "monetary policy rule set" and "industrial investment restriction rules") are matched from the financial regulatory rule library, and the subspace operation constraint group is loaded as the modulation parameter.
[0077] Targeted modulation of three types of rules: Macroeconomic trend subspace (risk matching rule): Use feature selection thresholds (such as 0.8 for conservative users) to modulate rule strictness: the original rule of "retain features with trend stability ≥ 0.6" is changed to "retain features with ≥ 0.8", filtering out more high-volatility macroeconomic indicators.
[0078] Industry-related subspace (target adaptation rule): Use node filtering conditions (such as "new energy + manufacturing") to expand the scope of the rule: the original rule only searches for the target company (such as CATL), and after the expansion, it also searches for its upstream lithium mining companies (such as Ganfeng Lithium), requiring a correlation degree of ≥0.7 (transaction proportion ≥70%).
[0079] Text signal pulse subspace (scale response rule): Use the sensitivity attenuation coefficient (such as 0.4 for 10 million funds) to scale the rule threshold: the original rule "text sentiment polarity ≤-0.5 triggers an alert" is changed to "≤-0.5×0.4=-0.2 triggers an alert", reducing the response of large amounts of funds to slightly negative text.
[0080] In the above step S53, macroeconomic trend subspace: the modulated risk matching rule is executed to extract segments that meet the threshold from the long-term trend characteristics (e.g., conservative users extract the trend segment of "GDP has grown for five consecutive years with a volatility of less than 3%").
[0081] Industry correlation subspace: Execute the expanded target adaptation rules to retrieve the industry chain paths that meet the filtering conditions (such as "new energy vehicle companies → battery suppliers → lithium mining companies", ensuring that the correlation degree of each node is ≥0.7).
[0082] Text signal pulse subspace: Executes scaled response rules after scaling to filter text event labels (e.g., 10 million yuan of funds filters negative labels such as "sentiment polarity ≤ -0.2 and related to new energy", such as "lithium ore prices plummet").
[0083] The trend segments (such as "macroeconomic stable growth"), industrial chain paths (such as "new energy core supply chain"), and text event tags (such as "short-term lithium price risk") are integrated according to weights: trend segment weight 0.5 (long-term basis), industrial chain path weight 0.3 (medium-term layout), and text tag weight 0.2 (short-term adjustment signal); generate personalized content assembly strategies, including asset allocation directions (such as "increase holdings of new energy midstream companies"), risk warnings (such as "be wary of lithium price fluctuations"), and adjustment trigger conditions (such as "reduce holdings when GDP growth rate is <5%").
[0084] In a preferred embodiment of the present invention, step S6 above: based on the personalized content assembly strategy and the risk calibration feature tensor, the neural symbolic rule engine performs rule-constrained content generation, and outputs an explainable investment report containing an investment portfolio plan, data traceability items, and dynamic adjustment instructions, including: Step S61: Parse the investment framework instructions, traceability path instructions, and adjustment logic tags in the personalized content assembly strategy, bind the three types of logic elements to the risk matching rules, traceability rules, and event response rules in the preset rule library, and generate a rule instantiation instruction set; Step S62: Based on the rule instantiation instruction set, a directed search is performed in the risk calibration feature tensor. In response to the risk matching rule, a return volatility evidence vector is extracted from the macroeconomic trend subspace; in response to the tracing rule, a financial report key index sequence is extracted from the industry association subspace; and in response to the event response rule, a policy-related signal label group is captured from the text signal subspace. Step S63: Input the rule instantiation instruction set, the return volatility evidence vector, the financial report key index sequence, and the policy-related signal label group into the neural symbolic rule engine, and execute the following: inject the return volatility evidence vector to generate an investment portfolio plan, associate the financial report key index sequence to generate a data traceability entry, and integrate the policy-related signal label group to generate a dynamic adjustment description; Step S64: Based on the investment portfolio plan, data traceability items and dynamic adjustment instructions, assemble a machine-verifiable investment report according to the financial regulatory template number.
[0085] In an embodiment of the present invention, core elements of the personalized content assembly strategy, such as the investment framework, traceability path, and adjustment logic, are bound to the risk matching, traceability, and event response rules in the preset rule library. This generates a rule-instantiated instruction set, transforming the abstract strategy logic into an executable "rule-operation" correspondence, thus avoiding logical deviations or omissions during content generation. Based on the instruction set, the return volatility evidence vector, financial report key index sequence, and policy-related signal label group are extracted from the risk calibration feature tensor, significantly enhancing the persuasiveness of the report. A neural symbolic rule engine is used to collaboratively generate investment portfolio plans, data traceability entries, and dynamic adjustment instructions. The investment portfolio plan is driven by the return volatility evidence vector, the data traceability entries are directly linked to the financial report index, and the dynamic adjustment instructions incorporate policy signals. Collaborative generation prevents disconnected content from different parts of the report, improving its integrity and internal relevance. Reports are assembled according to financial regulatory template numbers to ensure that the format and content meet regulatory requirements. Furthermore, the machine-verifiable nature of the report means that the portfolio logic, data traceability path, and adjustment basis in the report can all be verified through system backtesting.
[0086] In the embodiments of the present invention, when applied specifically, it can be achieved through the following technical solutions, for example: In step S61, the personalized content assembly strategy is structurally decomposed to extract three core logical elements: Investment framework instructions: These include asset allocation directions (e.g., "increase holdings in the new energy sector"), risk control requirements (e.g., "single asset share ≤ 30%)," and expected return ranges (e.g., "annualized return 5%-8%). Traceability path indication: specify the data source and search path (e.g., "Enterprise Financial Report - 2023Q3 - Net Profit", "Macroeconomic Database - GDP Growth Rate"); Adjustment logic tag: Define the trigger conditions for strategy adjustment (such as "reduce holdings when the proportion of negative sentiment signals in industry texts is greater than 60%" and "reduce stock positions when GDP growth rate is less than 5%").
[0087] Associate logical elements with preset rule bases: bind investment framework instructions to risk matching rules (such as "the proportion of conservative user stocks is ≤20%); bind traceability path instructions to traceability rules (such as "financial report data must be marked with the issuing agency, timestamp, and checksum"); adjust logical tags to bind event response rules (such as "execute the reduction rule when the proportion of negative emotional signals in the text exceeds the threshold").
[0088] Convert abstract rules into specific operational instructions (such as "according to the conservative risk matching rule, screen assets with volatility less than 5%; the traceability rule requires that all data be labeled with Wind codes; the event response rule trigger threshold is set to 60% of the text's negative sentiment signals").
[0089] In step S62, in the macroeconomic trend subspace, evidence is retrieved based on the risk matching rules of the rule-instantiated instruction set: when a conservative user instructs "volatility < 5%", the characteristic vectors corresponding to assets with return volatility < 5% in the past three years (such as the historical yield curves of government bonds and high-rated bonds) are screened; key parameters in the vector (such as maximum drawdown and Sharpe ratio) are extracted to form a return volatility evidence vector, which serves as the risk-return basis for the investment portfolio.
[0090] In the industry-related subspace, data indexes are extracted according to the traceability path instructions: for the "corporate financial report-2023Q3-net profit" path, the page number of the specific financial report document (such as page 8 of a company's 2023Q3 financial report), data verification code (such as hash value "a3b7c9..."), and release channel (such as the official website of the exchange) are located; the index information is arranged in chronological order to generate a financial report key index sequence to ensure that the data can be traced back to the original source.
[0091] In the text signal pulse subspace, relevant labels are filtered according to event response rules: when the instruction "the proportion of negative emotional signals in the text is greater than 60%, triggering adjustment" is given, text events related to investment directions (such as new energy) are retrieved, and labels with emotional polarity less than 0 (such as "lithium mining overcapacity" and "subsidy reduction rumors") are extracted; they are sorted by correlation strength (such as semantic similarity with the new energy industry ≥ 0.8) to generate a policy-related signal label group as a trigger signal for dynamic adjustment.
[0092] In the above step S63, the symbol rule layer constructs a framework: Investment portfolio framework: Based on risk matching rules, fill in the parameters of the investment framework instructions (e.g., a conservative user configuration of "40% government bonds, 30% high-grade bonds, 20% blue-chip stocks, 10% cash") to form a compliant plan framework; Data traceability entry framework: Using traceability rules, convert the financial report key index sequence into a traceability entry template (e.g., "Data source: XX Company 2023 Q3 Financial Report (page 8, checksum a3b7c9); Update time: 2023-10-15"); Dynamic adjustment description framework: Based on event response rules, the adjustment logic tags are converted into condition-action expressions (such as "IF the proportion of negative sentiment signals in new energy texts is greater than 60%, THEN reduce holdings of assets in this industry to 10%").
[0093] Neural network layer injection evidence: Inject the return volatility evidence vector into the skeleton, supplement specific targets and weights (such as "Increase holdings of government bonds: XX government bonds (volatility 2.3%, accounting for 40%); reduce holdings of high-volatility stocks: XX stocks (volatility 12%, excluded)"); add data summaries to the traceability entry framework (such as "XX company's net profit in Q3 2023 increased by 15% year-on-year"), and associate index sequences to ensure traceability; integrate policy-related signal label groups into expressions, and supplement logical explanations (such as "If the label 'lithium ore prices plummet' appears, it will trigger a reduction in holdings, because the label has a correlation of 0.9 with the cost sensitivity of the new energy industry chain").
[0094] In step S64 above, the investment portfolio plan, data traceability items, and dynamic adjustment instructions are divided into independent modules according to financial regulatory requirements: the investment portfolio module includes the asset list, weights, and risk ratings (corresponding to the regulatory template number "INV-001"); the traceability module includes data sources, indexes, and verification information (corresponding to the template "TRC-002"); and the adjustment instructions module includes trigger conditions, operation logic, and compliance basis (corresponding to the template "ADJ-003").
[0095] Standardize fonts, field order, and risk warnings (such as "past performance does not represent future returns") according to the template format; add machine-readable verification marks to each module: the portfolio module attaches the hash value of the asset allocation (to verify whether it has been tampered with); the traceability module embeds the digital signature of the data source (such as the official seal of the exchange); the entire report generates a unique verification code to support real-time verification by the regulatory system or third-party platform.
[0096] like Figure 2 As shown, the embodiment of the present invention also provides a large model content dynamic assembly system driven by multi-source heterogeneous domain data, including: The acquisition module is used to acquire structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment processing; The multi-source module is used to extract multimodal features based on the fused dataset through the spatiotemporal graph neural network, and output the multi-source feature embedding tensor through dynamic weight fusion; The calibration module is used to generate feature vectors by locating key nodes in the high-dimensional space of the multi-source feature embedding tensor. Specifically, it captures the long-term trend vectors of macroeconomic indicators in the structured feature aggregation area and the short-term mutation vectors of multi-source text sentiment signals in the unstructured feature aggregation area. Based on the feature vectors, it connects the long-term trend vectors and the short-term mutation vectors to generate a feature calibration axis. The semantic projection length of the feature calibration axis in the financial event knowledge graph is calculated. The risk calibration module is used to measure the spatial angle between the axis and the preset ideal decision plane based on the semantic projection length. When the spatial angle exceeds the dynamic threshold, a numerical correction is generated based on market volatility and user risk coefficient, and a multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor. A conversion module, which converts the risk calibration feature tensor into a personalized content assembly strategy based on the user profile; The engine module is used to assemble strategies and risk-calibrated feature tensors based on personalized content, perform rule-constrained content generation through a neural symbolic rule engine, and output an explainable investment report containing investment portfolio plans, data traceability entries, and dynamic adjustment instructions.
[0097] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A large model content dynamic assembly method driven by multi-source heterogeneous domain data, characterized by: The method comprises: Step S1: Obtain structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment processing; Step S2: Based on the fused dataset, multimodal features are extracted through the spatiotemporal graph neural network, and multi-source feature embedding tensors are output through dynamic weight fusion; Step S3: In the high-dimensional space of the multi-source feature embedding tensor, generate feature vectors by locating key nodes. That is, capture the long-term trend vectors of macroeconomic indicators in the structured feature aggregation area, and capture the short-term mutation vectors of multi-source text sentiment signals in the unstructured feature aggregation area. Based on the feature vectors, connect the long-term trend vectors and the short-term mutation vectors to generate a feature calibration axis. Calculate the semantic projection length of the feature calibration axis in the financial event knowledge graph. Step S4: The spatial angle between the axis and the preset ideal decision plane is measured based on the semantic projection length. When the spatial angle exceeds a dynamic threshold, a numerical correction is generated based on the market volatility and the user risk factor, and the multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor. Step S5: Convert the risk calibration feature tensor into a personalized content assembly strategy based on the user profile; Step S6: Based on the personalized content assembly strategy and risk calibration feature tensor, the neural symbolic rule engine is used to perform rule-constrained content generation, and an explainable investment report containing investment portfolio plans, data traceability items, and dynamic adjustment instructions is output.
2. The method for dynamically assembling large model content driven by multi-source heterogeneous domain data according to claim 1 is characterized in that: Step S1: Obtain structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment, including: Acquire structured, semi-structured, and unstructured raw data in the financial field to obtain original heterogeneous data sets; Based on the original heterogeneous data set, missing value filling and outlier smoothing are performed on structured data, table topology analysis and semantic conflict elimination are performed on semi-structured data, and noise filtering and sentiment calibration are performed on structured data to obtain a preliminary screening quality data set; Based on the initial screening quality dataset, we use the financial event knowledge graph to drive spatiotemporal alignment, aligning the timestamps of low-frequency economic indicators and high-frequency event data streams in the temporal dimension. In the spatial dimension, we map enterprise addresses and geographic tag data to a unified coordinate grid to obtain a spatiotemporal correlation dataset. Modal uniform conversion is performed on the spatiotemporal correlation dataset, that is, the numerical data is normalized to the interval [0, 1], the text data is encoded into a semantic vector of preset dimensions, and the tabular data is converted into a graph node embedding vector to generate a fused dataset.
3. The method for dynamically assembling large model content driven by multi-source heterogeneous domain data according to claim 2 is characterized in that: Step S2: Based on the fused dataset, multimodal features are extracted through the spatiotemporal graph neural network, and multi-source feature embedding tensors are output through dynamic weight fusion, including: The structured data in the fusion dataset is input into the time series convolution layer to capture the cyclical evolution of macroeconomic indicators and output the numerical evolution feature vector; The semi-structured data in the fusion dataset is input into the graph attention network to analyze the equity linkage chain and supply chain topology of corporate financial reports and industry reports, and output the inter-table linkage feature vector; The unstructured data in the fusion dataset is input into the multi-head text encoder, and the text sentiment extreme value is calculated based on the weighted financial sentiment dictionary, and the text semantic feature vector is output; The numerical evolution feature vector, inter-table association feature vector and text semantic feature vector are input into the dynamic weight allocation unit. By real-time scanning of the policy nodes and market fluctuation signals in the financial event knowledge graph, the preset weight allocation rules are matched according to the event type, and the numerical evolution feature vector, inter-table association feature vector and text semantic feature vector are weighted spliced and tensor regularized to output the multi-source feature embedding tensor.
4. The method for dynamically assembling large model content driven by multi-source heterogeneous domain data according to claim 3 is characterized in that: Step S3: In the high-dimensional space of the multi-source feature embedding tensor, feature vectors are generated by locating key nodes. That is, the long-term trend vectors of macroeconomic indicators are captured in the structured feature aggregation area, and the short-term mutation vectors of multi-source text sentiment signals are captured in the unstructured feature aggregation area. Based on the feature vector, the long-term trend vector and the short-term mutation vector are connected to generate the feature calibration axis; Calculate the semantic projection length of the feature calibration axis in the financial event knowledge graph, including: Based on the high-dimensional spatial geometric structure of the multi-source feature embedding tensor, the long-term trend vector of structured data is captured in the first feature area and the macroeconomic trend nodes are located. At the same time, the short-term mutation vector of unstructured data is captured in the second feature area and the mutation nodes of text sentiment signals are located. Connecting the macroeconomic trend node and the text sentiment signal mutation node to form a feature calibration axis; Calculate the direction vector modulus of the characteristic calibration axis as the axis reference length, and decompose the axis direction vector into policy dimension and market dimension components to generate a two-dimensional direction identifier; Using the two-dimensional direction identifier as the search key, the event node sequence matching the policy market label is retrieved in the financial event knowledge graph, the projection scalar of the feature calibration axis vector in the conduction chain direction of the event node sequence is calculated, and the accumulated value of the shortest path weight between the conduction chain nodes is taken as the semantic projection length.
5. The method for dynamically assembling large model content driven by multi-source heterogeneous domain data according to claim 4 is characterized in that: Step S4: The spatial angle between the axis and the preset ideal decision plane is measured based on the semantic projection length. When the spatial angle exceeds the dynamic threshold, a numerical correction is generated based on the market volatility and the user risk factor, and the multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor, including: Based on the semantic projection length and the real-time market volatility index, a dynamic deviation threshold is generated; and the spatial angle between the feature calibration axis and the preset rational decision plane is measured; When the spatial angle exceeds the dynamic deviation threshold, the angle difference of the spatial angle exceeding the threshold is calculated, the fluctuation tolerance coefficient in the user risk profile is taken, and the angle difference is multiplied by the fluctuation tolerance coefficient and then multiplied by the inverse of the semantic projection length to obtain a scalar correction value; The scalar correction amount is converted into a unit vector in the axis direction, and the characteristic component in the axis direction of the characteristic tensor is linearly enhanced or attenuated to obtain the corrected characteristic tensor; By calculating the variance value of each dimension of the feature tensor, scaling the standard deviation of the dimensions that exceed the overall variance mean, and constraining the value range of the scaled feature tensor to be between zero and one, a risk calibration feature tensor is generated.
6. The method for dynamically assembling large model content driven by multi-source heterogeneous domain data according to claim 5 is characterized in that: Step S5: Convert the risk calibration feature tensor into a personalized content assembly strategy based on the user profile, including: The risk calibration feature tensor is orthogonally projected and separated into a macroeconomic trend subspace, an industry association subspace, and a text signal pulse subspace. A subspace operation constraint group is generated based on user profiles, mapping the user risk level to a feature selection threshold in the trend subspace. Investment targets are converted into node filtering conditions in the industry subspace. Fund size is quantified as a sensitivity attenuation coefficient in the text signal subspace. In response to real-time market event signals, the financial regulatory rule base is activated using the subspace operation constraint group as a modulation parameter, and three types of rules are modulated: a risk matching rule acting on the macroeconomic trend subspace is modulated using a feature selection threshold, a target adaptation rule acting on the industry association subspace is expanded through a node filtering condition, and a scale response rule acting on the text signal impulse subspace is scaled using a sensitivity attenuation coefficient; The modulated risk matching rule is executed in the trend subspace to extract trend fragments; the expanded target adaptation rule is executed in the industry association subspace to retrieve the industrial chain path; the scaled scale response rule is executed in the text signal subspace to filter text event labels; and the trend fragments, industrial chain paths and text event labels are integrated to generate a personalized content assembly strategy.
7. The method for dynamically assembling large model content driven by multi-source heterogeneous domain data according to claim 6 is characterized in that: Step S6: Based on the personalized content assembly strategy and the risk calibration feature tensor, the neural symbolic rule engine performs rule-constrained content generation and outputs an explainable investment report containing investment portfolio solutions, data traceability items, and dynamic adjustment instructions, including: Parse the investment framework instructions, traceability path instructions, and adjustment logic tags in the personalized content assembly strategy, bind the three types of logical elements to the risk matching rules, traceability rules, and event response rules in the preset rule library, and generate a rule instantiation instruction set; Based on the rule instantiation instruction set, a directed search is performed in the risk calibration feature tensor. In response to the risk matching rule, the return volatility evidence vector is extracted from the macroeconomic trend subspace; in response to the tracing rule, the financial report key index sequence is extracted from the industry association subspace; in response to the event response rule, the policy association signal label group is captured from the text signal subspace; The rule instantiation instruction set, along with the return volatility evidence vector, financial report key index sequence, and policy-related signal label group, is fed into the neural symbolic rule engine. The engine then executes the following steps: inject the return volatility evidence vector to generate an investment portfolio plan, associate the financial report key index sequence to generate data provenance entries, and integrate the policy-related signal label group to generate dynamic adjustment instructions. Based on the investment portfolio plan, data traceability items and dynamic adjustment instructions, a machine-verifiable investment report is assembled according to the financial regulatory template number.
8. A large model content dynamic assembly system driven by multi-source heterogeneous domain data, the system implementing the method according to any one of claims 1 to 7, characterized in that: include: The acquisition module is used to acquire structured data, semi-structured data, and unstructured data in the financial field, and generate a fused data set through spatiotemporal alignment processing; The multi-source module is used to extract multimodal features based on the fused dataset through the spatiotemporal graph neural network, and output the multi-source feature embedding tensor through dynamic weight fusion; The calibration module is used to generate feature vectors by locating key nodes in the high-dimensional space of the multi-source feature embedding tensor. This module captures the long-term trend vectors of macroeconomic indicators in the structured feature aggregation area and the short-term mutation vectors of the multi-source text sentiment signals in the unstructured feature aggregation area. Based on the feature vector, the long-term trend vector and the short-term mutation vector are connected to generate the feature calibration axis; Calculate the semantic projection length of the feature calibration axis in the financial event knowledge graph; The risk calibration module is used to measure the spatial angle between the axis and the preset ideal decision plane based on the semantic projection length. When the spatial angle exceeds the dynamic threshold, a numerical correction is generated based on market volatility and user risk coefficient, and a multi-source feature embedding tensor is injected along the feature calibration axis to generate a risk calibration feature tensor. A conversion module, which converts the risk calibration feature tensor into a personalized content assembly strategy based on the user profile; The engine module is used to assemble strategies and risk-calibrated feature tensors based on personalized content, perform rule-constrained content generation through a neural symbolic rule engine, and output an explainable investment report containing investment portfolio plans, data traceability entries, and dynamic adjustment instructions.
9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Cited By
Dynamic data driven adaptive content generation and risk control method and system and medium
CN121030098A
Virtual model assembling method based on knowledge graph and mechanism embedding
CN121234777A