Financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment
By utilizing a financial data intelligent interaction system based on retrieval enhancement and multimodal alignment, and employing technologies such as adaptive vector generation models and ViT-L/14 models, the system addresses the problem of missing cross-modal semantic associations in financial data systems, achieving efficient and accurate multimodal data interaction.
Patent Information
- Application Number
- CN202511019733.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
AI Technical Summary
Existing intelligent interaction systems for financial data lack visual-language joint encoding capabilities, making it impossible to perform deep semantic understanding. This results in a lack of cross-modal semantic associations, low retrieval efficiency, and low accuracy when faced with complex semantic queries and structured data filtering.
A financial data intelligent interaction system based on retrieval enhancement and multimodal alignment is adopted, including a request parsing module, a multimodal vector generation module, a data filtering module, a cross-modal alignment module, a fusion generation module, and a compliance verification module. Through adaptive vector generation model, dual-path recall engine, ColBERT-style post-interaction architecture, ViT-L/14 model, and triple loss function, semantic matching and cross-modal alignment of multimodal data are achieved.
It improves the efficiency and accuracy of intelligent interaction of financial data, enhances the flexibility and precision of hybrid sorting, solves the problem of modal fragmentation of financial data, and achieves efficient and accurate cross-modal retrieval.
Smart Images

Figure CN120910230A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the cross field of financial technology and artificial intelligence, in particular to a financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment. BACKGROUND
[0002] At present, financial data is the digital trace of recording fund flow, asset status and market behavior, and its essence is the quantitative expression of the value exchange process. Financial data plays an important role in modern society, not only helping enterprises make wise decisions, but also having a profound impact on the entire financial industry. From investment analysis to risk management, from market prediction to economic policy, the application field of financial data is wide and diverse. Therefore, it is crucial to intelligently interact with financial data.
[0003] In the existing financial data intelligent interaction system, text reports, transaction charts (K-line charts / bar charts) and structured tables are stored independently, resulting in a lack of cross-modal semantic association. Traditional financial data intelligent interaction systems lack visual-linguistic joint coding capabilities, and existing financial data intelligent interaction systems fail to perform deep semantic understanding when performing vector retrieval, cannot improve semantic matching accuracy through dynamic context modeling, and the inverted index method based on keyword matching is difficult to handle complex semantic queries. Single vector retrieval has low accuracy when filtering structured data, resulting in low efficiency of financial data intelligent interaction based on retrieval enhancement and multi-modal alignment, and there is room for improvement. SUMMARY
[0004] In order to improve the efficiency of financial data intelligent interaction based on retrieval enhancement and multi-modal alignment, the present application provides a power information network fault positioning method.
[0005] The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment provided by the present application adopts the following technical solutions:
[0006] The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment comprises:
[0007] The request analysis module is configured to receive a user query request, analyze the semantic intent and time range limitation condition to obtain query text information;
[0008] The multi-modal vector generation module is configured to dynamically encode the query text based on the query text information, generate multiple feature semantic vectors and combine them to form a dynamic semantic vector library;
[0009] The data filtering module is configured to start a two-way recall engine and perform in parallel to obtain recall result information, wherein the recall result information includes a plurality of sub-data retrieval information, a dynamic weight value S is calculated by weighting each sub-data retrieval information, each sub-data retrieval information is sorted based on the dynamic weight value S, and a plurality of multi-modal data information is filtered and obtained;
[0010] The cross-modal alignment module is configured to perform feature matching on the multi-modal data information and the feature semantic vector, generate data comparison information, capture chart content in the multi-modal data information to obtain multi-modal data chart information, perform semantic space alignment on the query text information and the multi-modal data chart information, and output semantic space alignment information, and output a cross-modal alignment completion signal when receiving the data comparison information and the semantic space alignment information;
[0011] The fusion generation module is configured to fuse the multi-modal data information to generate a text analysis report, perform text traceability labeling and chart highlighting display analysis to obtain attached labeling information;
[0012] The compliance verification module is configured to verify whether the report content conforms to the specification by using a compliance checking engine, modify the content if it does not conform to the specification, and label the attached labeling information on the text analysis report to obtain a final version of the text analysis report if it conforms to the specification, and display the final version of the text analysis report on the user terminal.
[0013] Preferably, the natural language query in the financial field input by the user is obtained to obtain initial query information, and the initial query information is cached;
[0014] The initial query information is subjected to semantic intent analysis to obtain instruction semantic intent information, and the initial query information and the instruction semantic intent information are combined to form query instruction detailed information;
[0015] The initial query information is subjected to time range limitation to obtain instruction time limitation information;
[0016] The query instruction detailed information and the instruction time limitation information are combined to form the query text information.
[0017] Preferably, an adaptive vector generation model is created, the adaptive vector generation model is pre-trained based on a curriculum learning strategy, and a general knowledge base in the adaptive vector generation model is constructed and updated;
[0018] The adaptive vector generation model is preprocessed based on a convolutional neural network and updated, wherein a dynamic attention pooling layer is used to replace an average pooling layer in the convolutional neural network;
[0019] Injecting a field adaptation layer into the adaptive vector generation model, adjusting and updating the applicable field of the adaptive vector generation model, so that the adaptive vector generation model is adapted to the financial field;
[0020] Inputting the query text information into the adaptive vector generation model, and subdividing the query text information based on a Tokenizer word segmentation algorithm to obtain a plurality of word pieces;
[0021] For a plurality of word pieces, dynamically adjusting the resolution based on an adaptive grid division algorithm to obtain block granularity information;
[0022] Based on the block granularity information, block encoding each word piece information to obtain a plurality of dynamic encoding information;
[0023] Based on the Embedding technology, each dynamic encoding information is converted into a continuous vector to obtain a plurality of feature semantic vectors, and a dynamic semantic vector library is constructed based on the plurality of feature semantic vectors.
[0024] Preferably, a two-way recall engine is obtained and started, the two-way recall engine obtains a large amount of candidate data set, and the large amount of candidate data set includes a plurality of candidate document vectors;
[0025] Based on the ColBERT-style late interaction architecture, each candidate document vector in the large amount of candidate data set is retrieved, each candidate document vector in the large amount of candidate data set is matched with the dynamic semantic vector library, and the similarity is judged to obtain the matching similarity information of each candidate document vector. The matching similarity information of each candidate document vector is compared with the matching similarity threshold, and the candidate document vector with the matching similarity information greater than or equal to the matching similarity threshold is recalled to obtain semantic vector retrieval information;
[0026] The content of the numerical value in the query instruction detailed information is extracted and marked as numerical limit condition information, the instruction time limit information and the structural features of the numerical limit condition information in the query text information are extracted based on the eight-direction Sobel operator in the edge detection algorithm to obtain topological invariant feature information, and a B+ tree index is constructed based on the topological invariant feature information. The B+ tree index can support time condition and numerical condition query, and data screening is performed based on the B+ tree index to obtain structured feature retrieval information; this step is executed in parallel with the previous step;
[0027] The semantic vector retrieval information and the structured feature retrieval information are combined to form recall result information, and the recall result information includes a plurality of sub-data retrieval information.
[0028] Preferably, judging the similarity of each sub-data retrieval information and the query text information in the text obtains a text similarity sim text , judging the similarity of each sub-data retrieval information and the query text information in the image obtains an image similarity sim image , judging the matching degree of each sub-data retrieval information and the query text information in the structured field obtains a structured matching degree match struct ;
[0029] Based on the text similarity sim text , the image similarity sim image and the structured matching degree match struct , a design score function S=ɑ·sim text +β·sim image +γ·match struct in the hybrid sorting algorithm is used to perform a weighted scoring operation on each sub-data retrieval information to obtain a dynamic weight value S of each sub-data retrieval information, wherein ɑ, β, γ are all dynamic weight coefficients, and the dynamic weight coefficients are predicted by LSTM.
[0030] Based on the dynamic weight value S of each sub-data retrieval information, a weighted sorting is performed on each sub-data retrieval information to obtain a dynamic retrieval sorting result.
[0031] A preset number of sub-retrieval data information with high ranking in the dynamic retrieval sorting result is selected and marked as filtered data information, and the multiple filtered data information is combined to form multi-modal data information.
[0032] Preferably, each of the filtered data information is input into the adaptive vector generation model, and the adaptive vector generation model performs subdivision on each of the filtered data information based on the Tokenizer word segmentation algorithm to obtain multiple comparison word units.
[0033] For the multiple comparison word units, each filtered data information is divided into grid based on an adaptive grid division algorithm to obtain comparison block granularity information of each filtered data information.
[0034] Global feature information of each filtered data information is obtained based on a ViT-L / 14 model.
[0035] Based on the comparison block granularity information of each filtered data information, local feature information of each filtered data information is obtained by performing local feature extraction on different grid blocks of each filtered data information.
[0036] Multiple local feature information of each filtered data information is respectively calculated with multiple feature semantic vectors in the dynamic semantic vector library by MaxSim interaction to generate data comparison information.
[0037] Preferably, the semantic content of each filtered data information is captured based on the ViT-L / 14 model to obtain multi-modal data chart information;
[0038] According to the query text information and the multi-modal data chart information, the query text information and the multi-modal data chart information are aligned in the semantic space based on a triple loss function L = max(0, sim(q, d + )- sim(q, d - )+ margin), and semantic space alignment information is output, where q is the query text information, d + is the positive sample filtered data information, d - is randomly sampled irrelevant data, and margin is 0.2;
[0039] When the data comparison information is received and the semantic space alignment information is received, it is determined that each filtered data information and the query text information are aligned across modalities, and a cross-modal alignment completion signal is output.
[0040] Preferably, when the cross-modal alignment completion signal is received, a plurality of filtered data information in the multi-modal data information is fused to generate a text analysis report;
[0041] Based on the attention weight positioning mechanism, data tracing is performed on the text in the multi-modal fusion data to obtain data tracing mark information;
[0042] The intersection over union of the multi-modal data chart information and the feature semantic vector is calculated to obtain a feature intersection over union, and a region in the multi-modal data chart information where the feature intersection over union is greater than or equal to a preset target intersection over union is marked as visual highlight region mark information; it should be pointed out that the preset target intersection over union in the embodiment of the application is 0.65;
[0043] The data tracing mark information and the visual highlight region mark information are combined to form accompanying mark information.
[0044] Preferably, a rule verification engine is obtained, which is used to verify the text analysis report to determine whether it meets traceability requirements, and obtain a compliance verification result;
[0045] Based on the compliance verification result, if the traceability requirements are met, a compliance verification pass result is output, and if the traceability requirements are not met, the content of the text analysis report is modified until the traceability requirements are met and the compliance verification pass result is output;
[0046] When the compliance verification pass result is received, the data provenance mark information in the accompanying mark information and the visual highlight area are marked and displayed on the text analysis report to obtain a final version of the text analysis report;
[0047] The final version of the text analysis report is displayed on the user terminal.
[0048] In summary, the present application includes at least one of the following beneficial technical effects:
[0049] 1. The topological invariant features of the table data are calculated by the eight-direction Sobol operator to construct a B+ tree index that can support range queries, improving the retrieval efficiency;
[0050] 2. By using LSTM for prediction, the user's query text information is input into the model as a time series signal, thereby individually and instantaneously adjusting the weight in the hybrid sorting algorithm, greatly improving the flexibility and accuracy of hybrid sorting, and further improving the efficiency of intelligent interaction of financial data based on retrieval enhancement and multi-modal alignment;
[0051] 3. By MaxSim interactive calculation of each filtered data information in the multi-modal data information and multiple feature semantic vectors in the dynamic semantic vector library, the text and chart local and global matching pair generation data comparison information is improved, which improves the cross-modal retrieval accuracy, and the ViT-L / 14 model is used to capture the multi-modal data chart information, and the triple loss function L = max (0, sim (q, d + )-sim (q, d - )+margin) is used to realize the semantic space alignment of the query text information and the multi-modal data chart information, further improving the cross-modal retrieval accuracy, effectively solving the problem of financial data modal fragmentation, and further improving the efficiency of intelligent interaction of financial data based on retrieval enhancement and multi-modal alignment. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 The present embodiment mainly embodies the module schematic diagram of the financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment.
[0053] Reference signs: 1, request analysis module; 2, multi-modal vector generation module; 3, data filtering module; 4, cross-modal alignment module; 5, fusion generation module; 6, compliance verification module. DETAILED DESCRIPTION
[0054] The present application will be further described in detail below in conjunction with the accompanying drawings.
[0055] The present application discloses a financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment.
[0056] Referring to Figure 1 The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment comprises:
[0057] A request analysis module configured to receive a user query request, analyze semantic intention and time range limitation conditions to obtain query text information.
[0058] A multi-modal vector generation module configured to dynamically encode the query text based on the query text information, generate a plurality of feature semantic vectors and combine them to form a dynamic semantic vector library.
[0059] A data filtering module configured to start a double-path recall engine and perform in parallel to obtain recall result information, wherein the recall result information includes a plurality of sub-data retrieval information, the dynamic weight value S is obtained by weighting calculation on each sub-data retrieval information, and the multi-modal data information is obtained by sorting based on the dynamic weight value S.
[0060] A cross-modal alignment module configured to perform feature matching on the multi-modal data information and the feature semantic vector, generate data comparison information, capture the chart content in the multi-modal data information to obtain multi-modal data chart information, perform semantic space alignment on the query text information and the multi-modal data chart information, and output the semantic space alignment information, and output a cross-modal alignment completion signal when receiving the data comparison information and receiving the semantic space alignment information.
[0061] A fusion generation module configured to fuse the multi-modal data information to generate a text analysis report, perform text traceability labeling and chart highlighting display analysis to obtain attached labeling information.
[0062] A compliance verification module configured to verify whether the report content conforms to the specification by a compliance checking engine, modify the content if it does not conform to the specification, label the attached labeling information on the text analysis report if it conforms to the specification to obtain a final version of the text analysis report, and display the final version of the text analysis report on the user terminal.
[0063] The specific execution mode of the request analysis module is as follows:
[0064] Obtain the natural language query in the financial field input by the user to obtain initial query information, and perform cache processing on the initial query information.
[0065] Perform semantic intention analysis operation on the initial query information to obtain instruction semantic intention information, and combine the initial query information and the instruction semantic intention information to form query instruction detailed information.
[0066] Limit the time range of the initial query information to obtain instruction time limitation information.
[0067] The query instruction detailed information is combined with the instruction time limit information to form query text information.
[0068] The specific execution manner of the multi-modal vector generation module is as follows:
[0069] An adaptive vector generation model is created, the adaptive vector generation model is pre-trained based on a curriculum learning strategy, a general knowledge base in the adaptive vector generation model is constructed and updated.
[0070] The adaptive vector generation model is preprocessed and updated based on a convolutional neural network, and a dynamic attention pooling layer is used to replace an average pooling layer in the convolutional neural network.
[0071] It should be pointed out that the dynamic attention pooling layer is introduced to replace the average pooling layer in the embodiments of the present application, which reduces the vector similarity calculation error of the financial term by 18% and improves the accuracy of the feature semantic vector generation.
[0072] A domain adaptation layer is injected into the adaptive vector generation model, the applicable domain of the adaptive vector generation model is adjusted and updated, so that the adaptive vector generation model is adapted to the financial field.
[0073] It should be pointed out that in the embodiments of the present application, the domain adaptation layer (Domain Adaptation Layer) is injected into the adaptive vector generation model, which improves the vector representation accuracy of the financial term in the multi-modal vector generation, and improves the cosine similarity of professional concepts such as "illiquidity risk" by 23%, thereby improving the accuracy of the feature semantic vector generation.
[0074] The query text information is input into the adaptive vector generation model, and the adaptive vector generation model performs subdivision on the query text information based on the Tokenizer word segmentation algorithm to obtain a plurality of word units.
[0075] For the plurality of word units, the resolution is dynamically adjusted based on the adaptive grid division algorithm to obtain block granularity information.
[0076] Now an example is given for illustration: when the query text information is a financial chart, the block granularity information can be dynamically adjusted according to the complexity of the financial chart, wherein the block granularity interval threshold is 4*4 to 16*16. The complexity of the financial chart in the embodiments of the present application is general difficulty, and the block granularity information obtained is 8*8.
[0077] Based on the block granularity information, the word unit information is block encoded to obtain a plurality of dynamic encoding information.
[0078] Based on the Embedding technology, each dynamic encoding information is converted into a continuous vector to obtain a plurality of feature semantic vectors, and a dynamic semantic vector library is constructed based on the plurality of feature semantic vectors.
[0079] It should be pointed out that the Embedding technology in the embodiments of the present application refers to a core concept in machine learning and natural language processing, and its essence is a mathematical technology of converting discrete objects (such as words, sentences, images, etc.) into continuous vectors (real number arrays).
[0080] The specific execution mode of the data filtering module is as follows:
[0081] A two-way recall engine is obtained and started, the two-way recall engine obtains a massive candidate data set, and the massive candidate data set includes a plurality of candidate document vectors.
[0082] It should be pointed out that the two-way recall engine in the embodiments of the present application refers to a core architecture design in a recommendation system, an advertising system and a search system, and is mainly used for quickly screening out hundreds / thousands of related candidate objects (such as goods, articles, videos, etc.) in a massive candidate data set. Its core lies in recalling in parallel through two differentiated strategies, taking into account "efficiency + coverage" and "precision".
[0083] The ColBERT-style late interaction architecture is obtained and based on the ColBERT-style late interaction architecture, semantic vector retrieval is performed on each candidate document vector in the massive candidate data set, each candidate document vector in the massive candidate data set is matched with the dynamic semantic vector library, and similarity information of each candidate document vector is obtained. The matching similarity information of each candidate document vector is compared with the matching similarity threshold, and the candidate document vector with the matching similarity information greater than or equal to the matching similarity threshold is subjected to document recall to obtain semantic vector retrieval information.
[0084] The matching similarity threshold is 0.82. It should be pointed out that the ColBERT-style late interaction architecture in the embodiments of the present application allows fine-grained token-level matching of feature semantic vectors and candidate document vectors.
[0085] In actual application, on a similar data set of the business system itself, the Top-5 recall rate reaches 89.7%, while the recall rate of the traditional method is only 72.3%, and the response delay is reduced to 230 ms, while the benchmark system delay is 480 ms, meeting the real-time monitoring requirements and improving the Top-K document recall efficiency.
[0086] The content with a numerical value in the query instruction detailed information is extracted and marked as numerical value limiting condition information, and based on an eight-direction Sobel operator in an edge detection algorithm, instruction time limiting information and structural features of the numerical value limiting condition information in the query text information are extracted to obtain topological invariant feature information.
[0087] Specifically, in the embodiment of the application, the eight-direction Sobel operator is used to extract the structural features, and the convolution kernel weight matrix thereof is:
[0088]
[0089] Based on the topological invariant feature information, a B+ tree index is constructed, the B+ tree index can support time condition and numerical value condition queries, and data screening based on the B+ tree index obtains structured feature retrieval information. This step is executed in parallel with the previous step.
[0090] In actual application, the traditional vector retrieval does not fully utilize the deep semantic understanding ability of the adaptive vector generation model, and cannot improve the semantic matching precision through dynamic context modeling like the adaptive vector generation model. The inverted index method based on keyword matching is difficult to handle complex semantic queries such as "retail industry recovery trend after the epidemic", and the single vector retrieval has a 37% decrease in accuracy when facing structured data filtering (such as time range limitation). In the embodiment of the application, the topological invariant features of the table data are calculated by the eight-direction Sobol operator to construct a B+ tree index that can support range queries, thereby improving the retrieval efficiency.
[0091] The semantic vector retrieval information and the structured feature retrieval information are combined to form recall result information, and the recall result information includes a plurality of sub-data retrieval information.
[0092] The similarity of each sub-data retrieval information and the query text information in the text is judged to obtain the text similarity sim text The similarity of each sub-data retrieval information and the query text information in the image is judged to obtain the image similarity sim image The matching degree of each sub-data retrieval information and the query text information in the structured field is judged to obtain the structured matching degree match struct .
[0093] Based on the text similarity sim text , the image similarity sim image , and the structured matching degree match struct , a design score function S = a sim text + β sim image + γ match structThe weighting score operation is performed on each sub-data retrieval information to obtain a dynamic weight value S of each sub-data retrieval information, wherein a, β, and γ are all dynamic weight coefficients, and the dynamic weight coefficients are predicted by LSTM.
[0094] Specifically, [a, β, γ] = LSTM ([query_type, time_sensitivity, modal_bias]). The core idea of using LSTM to predict the dynamic weight coefficient in the embodiment of the application is to input the query text information of the user as a time sequence signal into the model, thereby performing individualized and instant adjustment on the weight in the hybrid ranking algorithm, and greatly improving the flexibility and accuracy of the hybrid ranking.
[0095] The dynamic retrieval ranking result is obtained by performing weighted ranking on each sub-data retrieval information based on the dynamic weight value S of each sub-data retrieval information.
[0096] A preset number of sub-retrieval data information with high ranking in the dynamic retrieval ranking result is selected and marked as filtered data information, and a plurality of filtered data information is combined to form multi-modal data information. In the embodiment of the application, the preset number is 10.
[0097] The specific execution mode of the cross-modal alignment module is as follows:
[0098] Each filtered data information is input into the adaptive vector generation model, and the adaptive vector generation model performs subdivision on each filtered data information to obtain a plurality of comparison word units based on the Tokenizer word segmentation algorithm.
[0099] For the plurality of comparison word units, each filtered data information is grid divided based on the adaptive grid division algorithm to obtain comparison block granularity information of each filtered data information.
[0100] Global feature information of each filtered data information is obtained based on the ViT-L / 14 model.
[0101] Based on the comparison block granularity information of each filtered data information, local feature information of each filtered data information is obtained by performing local feature extraction on different grid blocks of each filtered data information.
[0102] The plurality of local feature information of each filtered data information is respectively calculated with the plurality of feature semantic vectors in the dynamic semantic vector library by MaxSim interaction to generate data comparison information. It should be pointed out that the MaxSim interaction calculation in the embodiment of the application can use MACD column distribution index analysis calculation of K-line chart.
[0103] Specifically, when the filtered data information in the multi-modal data information is a financial chart and the block granularity information thereof is 8*8, the financial chart is divided into an 8*8 grid, each block is extracted by using the multi-granularity visual encoder of ColPali, global features are extracted by ViT-L / 14, and local feature extraction is performed using adaptive grid division, MaxSim interaction calculation is performed with the feature semantic vector in the shared space (such as the MACD columnar distribution of the K-line chart), and data comparison information is generated.
[0104] It should be pointed out that ColPali in the embodiments of the present application is a new visual retrieval model, which aims to realize efficient document retrieval through a visual language model. The core of the model is to use the ColBERT architecture and the PaliGemma model to combine visual information and text information to improve the accuracy and efficiency of retrieval.
[0105] Based on the ViT-L / 14 model, the semantics of the chart content in each filtered data information is captured to obtain multi-modal data chart information.
[0106] According to the query text information and the multi-modal data chart information, based on the triple loss function L = max(0, sim(q, d + )-sim(q, d - )+margin), the query text information and the multi-modal data chart information are aligned in the semantic space, and the semantic space alignment information is output, wherein q is the query text information, the positive sample d + is the filtered data information, the negative sample d - is randomly sampled irrelevant data, and margin is 0.2.
[0107] When the data comparison information is received and the semantic space alignment information is received, it is determined that each filtered data information and the query text information are aligned across modalities, and a cross-modal alignment completion signal is output.
[0108] It should be pointed out that the ViT-L / 14 model in the embodiments of the present application is a deep learning text encoding model based on the Vision Transformer (ViT) architecture, which is used to convert natural language text into vector representation for natural language understanding and semantic matching tasks.
[0109] In actual application, the traditional financial data system adopts independent storage of text reports, transaction charts (K-line chart / bar chart) and structured tables, resulting in lack of cross-modal semantic association. The traditional system lacks visual-linguistic joint coding capability and cannot capture visual semantic features in charts through multi-modal embedding like ColPali. For example, when the user queries "the relationship between Tesla's Q4 gross margin rate and stock price fluctuations", the system cannot automatically associate the data in the financial report text with the time series features of the stock trend chart.
[0110] In the present application, the MaxSim interaction calculation is performed between each filtered data information in the multi-modal data information and the plurality of feature semantic vectors in the dynamic semantic vector library, the data comparison information is generated by matching the local and global text and chart, the cross-modal retrieval accuracy is improved, the ViT-L / 14 model is used to capture the multi-modal data chart information, the triple loss function L = max (0, sim (q, d + )-sim (q, d - )+margin) is used to realize the semantic space alignment of the query text information and the multi-modal data chart information, further improve the cross-modal retrieval accuracy, and effectively solve the problem of financial data modal fragmentation.
[0111] The specific execution mode of the fusion generation module is as follows:
[0112] After receiving the cross-modal alignment completion signal, the plurality of filtered data information in the multi-modal data information is fused to generate a text analysis report.
[0113] Based on the attention weight positioning mechanism, data traceability information is obtained by performing data traceability on the text in the multi-modal fusion data.
[0114] The intersection-over-union of the multi-modal data chart information and the feature semantic vector is calculated to obtain a feature intersection-over-union, and the region with the feature intersection-over-union greater than or equal to a preset target intersection-over-union in the multi-modal data chart information is marked as visual highlight region marking information. It should be pointed out that the preset target intersection-over-union in the present application is 0.65.
[0115] The data traceability marking information and the visual highlight region marking information are combined to form the accompanying annotation information.
[0116] The specific execution mode of the compliance verification module is as follows:
[0117] The rule verification engine is obtained, which is used to verify the text analysis report and determine whether it meets the traceability requirement to obtain a compliance verification result.
[0118] Based on the compliance verification result, if the traceability requirement is met, a compliance verification pass result is output, and if the traceability requirement is not met, the content of the text analysis report is modified until the traceability requirement is met, and then the compliance verification pass result is output.
[0119] When the compliance verification pass result is received, the data source mark information in the accompanying mark information and the visual highlight area are marked and displayed on the text analysis report, and a final version of the text analysis report is obtained.
[0120] The final version of the text analysis report is displayed on the user terminal.
[0121] The above are preferred embodiments of the present application, which do not limit the protection scope of the present application, and therefore: any equivalent changes made on the basis of the structure, shape, principle of the present application should be covered within the protection scope of the present application.
Claims
1. A system for intelligent interaction with financial data based on search augmentation with multi-modal alignment, characterized in that, Comprise: Request analysis module, configured to receive user query request, parse semantic intent and time range limit condition to obtain query text information; Multi-modal vector generation module, configured to dynamically encode the query text based on the query text information, generate multiple feature semantic vectors and combine to form a dynamic semantic vector library; Data filtering module, configured to start a double-path recall engine and execute in parallel to obtain recall result information, which includes multiple sub-data retrieval information, and to calculate the dynamic weight value S by weighting each sub-data retrieval information, sort and filter the multi-modal data information based on the dynamic weight value S; Cross-modal alignment module, configured to perform feature matching on multi-modal data information and feature semantic vectors, generate data comparison information, capture chart content in multi-modal data information to obtain multi-modal data chart information, align the query text information and the multi-modal data chart information in semantic space, and output semantic space alignment information, and output cross-modal alignment completion signal when receiving data comparison information and receiving semantic space alignment information; Fusion generation module, configured to fuse multi-modal data information to generate a text analysis report, perform text tracing annotation and chart highlighting analysis to obtain attached annotation information; Compliance verification module, configured to verify whether the report content conforms to the specification by a compliance checking engine, and if not, modify the content, if so, mark the attached annotation information on the text analysis report to obtain a final version of the text analysis report, and display the final version of the text analysis report on the user terminal. 2.The financial data intelligent interaction system based on retrieval augmentation and multi-modal alignment of claim 1, wherein, The specific execution mode of the request analysis module is as follows: Obtain the user's natural language query in the financial field to obtain initial query information, and perform cache processing on the initial query information; Perform semantic intent analysis on the initial query information to obtain instruction semantic intent information, and combine the initial query information and the instruction semantic intent information to form query instruction detailed information; Limit the time range of the initial query information to obtain instruction time limit information; The query instruction detailed information and the instruction time limit information are combined to form the query text information. 3.The financial data intelligent interaction system based on retrieval augmentation and multi-modal alignment of claim 2, wherein, The specific execution mode of the multi-modal vector generation module is as follows: Create an adaptive vector generation model, pre-train the adaptive vector generation model based on a curriculum learning strategy, build and update the general knowledge base in the adaptive vector generation model; Preprocess the adaptive vector generation model based on a convolutional neural network and update the adaptive vector generation model, wherein a dynamic attention pooling layer is used to replace the average pooling layer in the convolutional neural network; Inject a domain adaptation layer into the adaptive vector generation model to adjust and update the applicable domain of the adaptive vector generation model, so that the adaptive vector generation model is adapted to the financial field; Input the query text information into the adaptive vector generation model, and the adaptive vector generation model subdivides the query text information into multiple word units based on the Tokenizer word segmentation algorithm; For multiple word units, the resolution is dynamically adjusted based on an adaptive grid division algorithm to obtain block granularity information; Based on the block granularity information, each word unit information is block encoded to obtain multiple dynamic encoding information; Based on the Embedding technology, each dynamic encoding information is converted into a continuous vector to obtain multiple feature semantic vectors, and a dynamic semantic vector library is constructed based on the multiple feature semantic vectors.
4. The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment according to claim 3, characterized in that, The specific execution mode of the data filtering module is as follows: A two-way recall engine is obtained and started, which obtains a large amount of candidate data set, and the large amount of candidate data set includes multiple candidate document vectors; Based on the ColBERT-style late interaction architecture, semantic vector retrieval is performed on each candidate document vector in the large amount of candidate data set, each candidate document vector in the large amount of candidate data set is matched with the dynamic semantic vector library, and the similarity is judged to obtain matching similarity information of each candidate document vector. The matching similarity information of each candidate document vector is compared with the matching similarity threshold, and the candidate document vector with matching similarity information greater than or equal to the matching similarity threshold is subjected to document recall to obtain semantic vector retrieval information; The content of the numerical value in the query instruction detail information is extracted and marked as numerical limit condition information. Based on the eight-direction Sobel operator in the edge detection algorithm, the instruction time limit information and the structural features of the numerical limit condition information in the query text information are extracted to obtain topological invariant feature information. Based on the topological invariant feature information, a B+ tree index is constructed, which can support time condition and numerical condition queries. Based on the B+ tree index, data screening is performed to obtain structured feature retrieval information. This step is executed in parallel with the previous step. The semantic vector retrieval information and the structured feature retrieval information are combined to form recall result information, which includes multiple sub-data retrieval information.
5. The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment according to claim 4, characterized in that, The specific execution mode of the data filtering module further includes: judging similarity of each sub-data retrieval information and the query text information in text to obtain text similarity sim text judging similarity of each sub-data retrieval information and the query text information in image to obtain image similarity sim image judging matching degree of each sub-data retrieval information and the query text information in structured field to obtain structured matching degree match struct Based on the text similarity sim text , the image similarity sim image , and the structured matching degree match struct , a design score function S = a • sim text + β • sim image + γ • match struct in a hybrid sorting algorithm is adopted to perform a weighted score operation on each sub-data search information to obtain a dynamic weight value S of each sub-data search information, wherein a, β, and γ are all dynamic weight coefficients, and the dynamic weight coefficients are predicted by an LSTM. Based on the dynamic weight value S of each sub-data retrieval information, each sub-data retrieval information is weighted and sorted to obtain a dynamic retrieval sorting result; Select the top pre-set number of sub-retrieval data information in the dynamic retrieval sorting result and mark it as filtered data information. The multiple filtered data information is combined to form multi-modal data information.
6. The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment according to claim 5, characterized in that, The specific execution mode of the cross-modal alignment module is as follows: Each of the filtered data information is input into the adaptive vector generation model, and the adaptive vector generation model subdivides each of the filtered data information based on the Tokenizer word segmentation algorithm to obtain multiple comparison word units; For multiple comparison word units, each filtered data information is grid divided based on an adaptive grid division algorithm to obtain comparison block granularity information of each filtered data information; Based on the ViT-L / 14 model, global feature information of each filtered data information is obtained through global feature extraction; Based on the comparison block granularity information of each filtered data information, local feature information of each filtered data information is obtained through local feature extraction of different grid blocks of each filtered data information; The MaxSim interaction calculation is performed between each local feature information of the filtered data information and the plurality of feature semantic vectors in the dynamic semantic vector library, and data comparison information is generated.
7. The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment according to claim 6, characterized in that, The specific execution mode of the cross-modal alignment module further includes: The semantic of the chart content in each filtered data information is captured based on the ViT-L / 14 model to obtain the multi-modal data chart information. According to the query text information and the multi-modal data chart information, the query text information is aligned with the multi-modal data chart information in a semantic space based on a triple loss function L = max(0, sim(q, d + ) - sim(q, d - ) + margin), and semantic space alignment information is output, wherein q is the query text information, a positive sample d + is filtered data information, a negative sample d - is randomly sampled irrelevant data, and margin is 0.2; When the data comparison information is received and the semantic space alignment information is received, it is determined that the cross-modal alignment of each filtered data information and the query text information is achieved, and a cross-modal alignment completion signal is output.
8. The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment according to claim 7, characterized in that, The specific execution mode of the fusion generation module is as follows: When the cross-modal alignment completion signal is received, the plurality of filtered data information in the multi-modal data information is fused to generate a text analysis report. Based on the attention weight positioning mechanism, data tracing is performed on the text in the multi-modal fusion data to obtain data tracing mark information. The intersection-over-union of the multi-modal data chart information and the feature semantic vector is calculated to obtain a feature intersection-over-union, and the region in the multi-modal data chart information with a feature intersection-over-union greater than or equal to a preset target intersection-over-union is marked as visual highlight region mark information; it should be pointed out that the preset target intersection-over-union in the embodiment of the application is 0.65; The data tracing mark information and the visual highlight region mark information are combined to form attached mark information.
9. The financial data intelligent interaction system based on retrieval enhancement and multi-modal alignment according to claim 8, characterized in that, The specific execution mode of the compliance verification module is as follows: A rule verification engine is obtained, which is used to verify the text analysis report to determine whether it meets the traceability requirement and obtain a compliance verification result; Based on the compliance verification result, if the traceability requirement is met, a compliance verification pass result is output, and if the traceability requirement is not met, the content of the text analysis report is modified until the traceability requirement is met, and then the compliance verification pass result is output; When the compliance verification pass result is received, the data tracing mark information in the attached mark information and the visual highlight region in the text analysis report are marked and displayed to obtain a final version of the text analysis report; The final version of the text analysis report is displayed on the user terminal.
Citation Information
Cited By
Cross-modal heterogeneous data retrieval method and system based on semantic information
CN121167001A
Intelligent power document generation method based on multi-modal memory fusion
CN121457451A