An intelligent data analysis and knowledge question answering method and system based on multimodal large model

By adopting multimodal large models, manifold learning enhancement algorithms, multi-task learning technology and recurrent neural network algorithms in the intelligent question-and-answer system, the problem of insufficient accuracy in existing systems when dealing with complex multimodal queries is solved, and more efficient and accurate data analysis and knowledge question-and-answer are achieved.

CN119577116BActive Publication Date: 2025-05-13MAXROCKY(BEIJING) INFO-TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510112720.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

When handling complex multimodal queries, existing intelligent question-and-answer systems lack accuracy, are difficult to capture the intrinsic relationships between different data forms, and have limited recursive processing capabilities, so they cannot fully simulate the dynamic time changes of sequence information.

Method used

Using intelligent data analysis and knowledge question-and-answer methods based on multimodal large models, the information integration of multimodal data is optimized through manifold learning enhancement algorithm and multi-task learning technology to generate a cross-modal unified representation. 同时,使用循环神经网络算法进行递归处理,捕捉序列信息间的依赖关系,并引入情境感知反馈机制,动态调整回答策略。

Benefits of technology

It improves the accuracy and response quality of intelligent data analysis and knowledge Q&A systems, enhances the robustness of multimodal data processing and the generalization ability of model, and can better understand and respond to complex multimodal queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577116B_ABST
    Figure CN119577116B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for intelligent data analysis and knowledge question and answer based on a multimodal large model. Among them, a multimodal query request is received, semantic parsing and intent recognition technology is performed, the core of the problem and the required knowledge field are determined, and a personalized problem understanding framework is generated; the manifold learning enhancement algorithm is used to perform manifold embedding processing to reveal the internal geometric structure, and multi-task learning technology is used to build a shared representation space, and multiple subtasks are trained at the same time to generate a cross-modal unified representation; a recurrent neural network algorithm is used for recursive processing, the internal state is recursively updated, and a hierarchical clustering method is used to build an organized information structure to generate an ordered information structure diagram; a situational awareness feedback mechanism is introduced to monitor instant reactions and environmental changes in real time, dynamically adjust the answer strategy, record instant feedback, and generate intelligent data analysis and knowledge question and answer solutions. The technical solution provided by this application improves the accuracy and user experience of intelligent data analysis and knowledge question and answer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence equipment technology, and in particular to an intelligent data analysis and knowledge question-answering method and system based on a multimodal large model. Background Art

[0002] With the explosive growth of information and the diversification of user needs, user query requests often contain multiple data forms, such as text, images, audio and video. This requires the system to not only receive and parse multimodal query requests from users, but also to perform semantic analysis and intent recognition on these requests to determine the core of the problem and the required knowledge areas. Generating a personalized understanding framework is the basis for ensuring accurate answers to user queries. In addition, the system needs to have the ability to efficiently process and integrate multimodal data in order to provide users with comprehensive and accurate answers.

[0003] Currently, many intelligent question-answering systems rely on a single-modal data processing method, that is, they mainly analyze text data. For complex queries involving multimodal data, existing systems usually use a discrete processing approach, analyzing different types of input data (such as text, images) independently, and then trying to merge the results. Although this method can meet basic needs to a certain extent, its effect is often unsatisfactory when faced with complex multimodal queries. In addition, some advanced systems have begun to introduce manifold learning enhancement algorithms and multi-task learning techniques, trying to optimize the information integration of multimodal data and improve transmission efficiency by building a shared representation space.

[0004] Although existing solutions have made some progress in some aspects, they still have significant shortcomings. First, due to the lack of an effective multimodal data integration mechanism, existing systems find it difficult to capture the inherent connections between different data forms, resulting in insufficient accuracy in intelligent data analysis and knowledge question answering. Second, the recursive processing capabilities of existing systems are limited and cannot fully simulate the dynamic time changes of sequence information, thus affecting the capture of complex dependencies. Finally, most systems fail to monitor users' immediate reactions and environmental changes in real time, and lack the ability to dynamically adjust answer strategies, making it difficult to ensure the relevance and accuracy of the answer content. Summary of the invention

[0005] The embodiments of the present application provide a method and system for intelligent data analysis and knowledge question answering based on a multimodal large model, so as to solve the problem of insufficient accuracy of intelligent data analysis and knowledge question answering in the prior art.

[0006] In a first aspect, the present application provides an intelligent data analysis and knowledge question answering method based on a multimodal large model, including:

[0007] Receive and parse multimodal query requests from users, perform semantic parsing and intent recognition technology on the multimodal query requests, determine the core of the question and the required knowledge domain in the multimodal query requests, and generate a personalized question understanding framework; the multimodal query requests contain a combination of data in multiple forms such as text, images, audio and video;

[0008] Using a manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is subjected to manifold embedding processing to reveal the intrinsic geometric structure of the multimodal data so as to optimize the information integration between the modalities in the multimodal data. A shared representation space is constructed using multi-task learning technology, and multiple subtasks are trained simultaneously in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a unified cross-modal representation.

[0009] Using a recurrent neural network algorithm, recursively processing the sequence information in the cross-modal unified representation, recursively updating the internal state of the sequence information, simulating the dynamic time changes of the sequence information to capture the dependencies between the sequence information, using a hierarchical clustering method, constructing an organized information structure for the sequence information through hierarchical aggregation, and gradually merging the closest data points in the sequence information according to the similarity measure in the organized information structure, providing deep knowledge mining, and generating an ordered information structure diagram;

[0010] A context-aware feedback mechanism is introduced to monitor users' immediate reactions and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record users' immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

[0011] Optionally, the manifold learning enhancement algorithm is used to perform manifold embedding processing on the multimodal data in the personalized question understanding framework to reveal the intrinsic geometric structure of the multimodal data to optimize the information integration between the modalities in the multimodal data, and a shared representation space is constructed using a multi-task learning technology. Multiple subtasks are trained simultaneously in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a cross-modal unified representation, including:

[0012] Extracting multimodal data features in the personalized question understanding framework to analyze potential associations between different data features in the multimodal data features and generate a multimodal association map;

[0013] Using a manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is processed into the multimodal association graph by manifold embedding, revealing the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding representation;

[0014] A shared representation space is constructed using multi-task learning technology, multiple interrelated subtasks are designed for the manifold embedding representation, and the interrelated subtasks are simultaneously trained in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a multi-task collaborative model;

[0015] The output results of each subtask in the multi-task collaborative model are obtained, mapped back to the shared representation space, and an adaptive regularization mechanism is introduced to dynamically perform regularization processing according to the output results of each subtask, so as to improve the generalization ability of the multi-task collaborative model and generate a cross-modal unified representation.

[0016] Optionally, the use of a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework into the multimodal association graph, revealing the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding representation, includes:

[0017] A similarity matrix is ​​introduced to evaluate the multimodal correlation in the multimodal association map, and a highly consistent part in the multimodal association map is identified. A local linear embedding technique is used to retain the local neighborhood structure during the identification process to generate a consistent local neighborhood structure.

[0018] Using a manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is subjected to manifold embedding processing into the multimodal association graph, and the consistent local neighborhood structure is used as a structural basis to reveal the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding intermediate result;

[0019] A global information fusion mechanism is introduced for the manifold embedding intermediate result, and the manifold embedding intermediate result is further optimized by combining the global context information in the personalized problem understanding framework to generate a global optimized embedding graph;

[0020] A weighted average strategy is introduced to assign different weights according to the importance of different data points in the global optimization embedding graph. A dimension selection technique is applied to select the dimension combination that is most representative of the multimodal data characteristics from the candidate dimensions generated by the global optimization embedding graph to generate a manifold embedding representation.

[0021] Optionally, the multi-task learning technology is used to construct a shared representation space, a plurality of interrelated subtasks are designed for the manifold embedding representation, and the interrelated subtasks are simultaneously trained in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a multi-task collaborative model, including:

[0022] identifying natural clusters in the manifold embedding representation by a clustering method to obtain concentrated areas in the manifold embedding representation, evaluating the internal consistency and external separability of each cluster in the natural clusters, and generating a manifold embedding cluster analysis result;

[0023] Adopting multi-task learning technology, constructing a shared representation space according to the manifold embedding cluster analysis result, designing multiple interrelated subtasks for the manifold embedding representation, and simultaneously training the interrelated subtasks in the shared representation space to improve the transfer efficiency of the personalized question understanding framework and generate a multi-task learning architecture;

[0024] For the multi-task learning framework, a task relevance matrix is ​​introduced to quantify the task relevance in the multi-task learning framework, so as to enhance the synergy between subtasks in the multi-task learning framework and generate a task synergy optimization graph;

[0025] In the shared representation space, an initial weight is assigned to each collaborative task in the task collaborative optimization graph, the initial weight is dynamically adjusted through a back-propagation algorithm, and an early stopping mechanism is introduced to prevent overfitting during the dynamic adjustment process, thereby generating a multi-task collaborative model.

[0026] Optionally, the recurrent neural network algorithm is used to recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, simulate the dynamic time change of the sequence information, so as to capture the dependency relationship between the sequence information, adopt a hierarchical clustering method, construct an organized information structure for the sequence information through a hierarchical aggregation method, and gradually merge the closest data points in the sequence information according to the similarity measurement in the organized information structure, provide deep knowledge mining, and generate an ordered information structure diagram, including:

[0027] Applying time series feature extraction technology to extract time-dependent components in the cross-modal unified representation, analyzing the time pattern and periodicity features in the time-dependent components, and generating a time series feature set;

[0028] Using a recurrent neural network algorithm, recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, input the internal state into the time series feature set, simulate the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, and generate a time series internal state diagram;

[0029] Adopting a hierarchical clustering method, constructing an organized information structure for the sequence information and the internal state diagram of the time series by hierarchical aggregation, gradually merging the closest data points in the sequence information according to the similarity measurement in the organized information structure, providing deep knowledge mining, and generating a hierarchical clustering structure diagram;

[0030] Identify the key nodes and paths in the hierarchical clustering structure diagram, apply a graph traversal algorithm, explore layer by layer from the root node to the leaf node in the key node, record the important transmission paths of ordered information in the layer-by-layer exploration process, and generate an ordered information structure diagram.

[0031] Optionally, the using of a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, input the internal state into the time series feature set, simulate the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, and generate a time series internal state diagram, including:

[0032] Applying a sliding window mechanism, extracting local time series features from the time series feature set, combining time-frequency transformation technology, capturing global periodic patterns based on the local time series features, and generating a refined time series feature set;

[0033] Using a recurrent neural network algorithm, recursively processing the sequence information in the cross-modal unified representation and the refined time series feature set, recursively updating the internal state of the sequence information, inputting the internal state into the time series feature set, simulating the dynamic time change of the sequence information, so as to capture the dependency relationship between the sequence information, and generating a recursively updated state record;

[0034] Mapping the internal state of each time step in the recursively updated state record to the corresponding time point, and applying an anomaly detection algorithm to identify and mark potential abnormal mapping transitions to generate a time series state mapping table;

[0035] The internal state information of all time steps in the time series state mapping table is integrated, and the mapping results in the time series state mapping table are converted into intuitive graphic representations by applying visualization technology to generate a time series internal state diagram.

[0036] Optionally, the context-aware feedback mechanism is introduced to monitor the user's immediate reaction and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question answering solutions, including:

[0037] Analyze the relationship between nodes and edges in the ordered information structure graph to identify the logical connection between different information elements in the ordered information structure graph, extract the information units and interaction modes represented by each node in the ordered information structure graph through pattern recognition technology, and generate a node and edge information table;

[0038] Introducing a context-aware feedback mechanism to capture the user's immediate reaction and changes in the surrounding environment, and combining the node and edge information table to convert the immediate reaction and changes in the surrounding environment into computable feature vectors to generate a computable feature vector group;

[0039] Dynamically adjust the answer strategy for the personalized question understanding framework, dynamically weight the computable feature vector group by performing context-dependent dynamic weight assignment on the ordered information structure graph, and generate a dynamic weighted feature vector;

[0040] Evaluate the relevance and accuracy of the answer content in the answer strategy, optimize the evaluation result in combination with the dynamic weighted feature vector, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

[0041] In a second aspect, the present application embodiment provides an intelligent data analysis and knowledge question answering system based on a multimodal large model, including:

[0042] A receiving module is used to receive and parse a multimodal query request from a user, perform semantic parsing and intent recognition technology on the multimodal query request, determine the core of the question and the required knowledge domain in the multimodal query request, and generate a personalized question understanding framework; the multimodal query request includes a combination of data in multiple forms such as text, image, audio and video;

[0043] A processing module, used to use a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework, reveal the intrinsic geometric structure of the multimodal data, optimize the information integration between the modes in the multimodal data, use multi-task learning technology to construct a shared representation space, and simultaneously train multiple subtasks in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a cross-modal unified representation;

[0044] A simulation module, for applying a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation, recursively updating the internal state of the sequence information, simulating the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, using a hierarchical clustering method to construct an organized information structure for the sequence information through a hierarchical aggregation method, and gradually merging the closest data points in the sequence information according to the similarity measurement in the organized information structure, providing deep knowledge mining, and generating an ordered information structure diagram;

[0045] The monitoring module is used to introduce a context-aware feedback mechanism, monitor the user's immediate reaction and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

[0046] In a third aspect, an embodiment of the present application provides a computing device, comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an intelligent data analysis and knowledge question-answering method based on a multimodal large model as described in the first aspect.

[0047] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements an intelligent data analysis and knowledge question-answering method based on a multimodal large model as described in the first aspect.

[0048] In an embodiment of the present application, a multimodal query request from a user is received and parsed, and semantic parsing and intent recognition technology are performed on the multimodal query request to determine the core of the problem and the required knowledge field in the multimodal query request, and generate a personalized problem understanding framework; the multimodal query request includes a combination of data in multiple forms such as text, image, audio and video; a manifold learning enhancement algorithm is used to perform manifold embedding processing on the multimodal data in the personalized problem understanding framework to reveal the intrinsic geometric structure of the multimodal data to optimize the information integration between the modes in the multimodal data, and a shared representation space is constructed using multi-task learning technology, and multiple subtasks are trained simultaneously in the shared representation space to improve the transmission efficiency of the personalized problem understanding framework and generate a cross-modal unified representation; a recurrent neural network algorithm is used to embed the multimodal data in the cross-modal unified representation The sequence information of the sequence information is recursively processed, the internal state of the sequence information is recursively updated, and the dynamic time change of the sequence information is simulated to capture the dependency relationship between the sequence information. The hierarchical clustering method is used to construct an organized information structure for the sequence information through hierarchical aggregation. According to the similarity measurement in the organized information structure, the closest data points in the sequence information are gradually merged to provide deep knowledge mining and generate an ordered information structure diagram; the context-aware feedback mechanism is introduced to monitor the user's immediate response and environmental changes in real time, and the answer strategy for the personalized question understanding framework is dynamically adjusted. By performing context-related dynamic weight allocation on the ordered information structure diagram, the relevance and accuracy of the answer content in the answer strategy are evaluated and optimized, and the user's immediate feedback on the answer content is recorded to generate intelligent data analysis and knowledge question answering solutions. By receiving and parsing multimodal query requests, a personalized question understanding framework is generated in combination with semantic parsing and intent recognition technology. This process can not only accurately capture the user's complex query needs, but also efficiently integrate across multiple data forms such as text, images, audio and video, thereby providing more accurate knowledge services. Furthermore, the manifold learning enhancement algorithm and recurrent neural network algorithm are used to optimize the information integration of multimodal data, and the answer strategy is dynamically adjusted through the context-aware feedback mechanism to ensure the relevance and accuracy of the answer content. This method greatly improves the interactive experience and response quality of the intelligent data analysis and knowledge question-answering system.

[0049] Furthermore, by constructing a shared representation space and training multiple subtasks simultaneously, the transfer efficiency of the personalized question understanding framework is improved, and finally a cross-modal unified representation is generated. This method not only enhances the robustness of multimodal data processing, but also improves the generalization ability of the model, enabling the system to maintain high performance in different application scenarios.

[0050] Furthermore, the internal state diagram of the time series is generated through time series feature extraction technology and recurrent neural network algorithm, and then the hierarchical clustering method is used to construct an organized information structure, gradually merging the closest data points to provide deep knowledge mining. Finally, the important transmission paths of ordered information are recorded through the graph traversal algorithm to generate an ordered information structure diagram. This method not only effectively captures the complex dependencies between sequence information, but also provides a structured foundation for subsequent knowledge mining, greatly improving the depth and accuracy of the analysis results.

[0051] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 A flowchart of an intelligent data analysis and knowledge question answering method based on a multimodal large model provided in an embodiment of the present application;

[0054] Figure 2 A schematic diagram of the structure of an intelligent data analysis and knowledge question answering system based on a multimodal large model provided in an embodiment of the present application;

[0055] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0057] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.

[0058] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0059] Figure 1 A flowchart of an intelligent data analysis and knowledge question answering method based on a multimodal large model is provided for an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0060] 101. Receive and parse a multimodal query request from a user, perform semantic parsing and intent recognition technology on the multimodal query request, determine the core of the question and the required knowledge domain in the multimodal query request, and generate a personalized question understanding framework; the multimodal query request includes a combination of data in multiple forms such as text, image, audio, and video;

[0061] In this step, multimodal query requests refer to queries raised by users in multiple data forms (such as text, images, audio, and video). These data forms can be used alone or in combination. Each data form carries different types of information and needs to be processed comprehensively to understand the user's intentions.

[0062] Semantic parsing and intent recognition technology aims to analyze the text content input by users through natural language processing technology and machine learning models to determine its semantic meaning and potential intentions. This technology can convert unstructured user queries into structured information for further processing.

[0063] The core of the problem and the required knowledge areas refer to the key problem points and related knowledge areas extracted from multimodal query requests, which helps the system focus on what users really care about and provide them with targeted answers.

[0064] The personalized question understanding framework is a structured representation generated by in-depth analysis of user queries, which contains user intent, the core of the question, and relevant background information.

[0065] In the embodiment of the present application, first, the system receives a multimodal query request from a user; second, the query request is analyzed using semantic parsing and intent recognition technology to determine the core of the problem and the required knowledge areas; third, a personalized understanding framework is generated based on the analysis results, which integrates data features of all modalities; finally, this framework serves as the basis for subsequent steps to ensure that the system can accurately understand and respond to user queries.

[0066] Suppose a user uploads a travel log containing text descriptions and pictures, asking about the historical background of a certain scenic spot. The system first receives and parses this multimodal query request, including the text description and picture content in the text; secondly, it uses semantic parsing and intent recognition technology to analyze the text and determine that the core of the user's question is "the historical background of a certain scenic spot"; thirdly, it combines the visual information in the picture to generate a personalized question understanding framework containing text and image features; finally, this framework will guide the system to search for relevant information in the next step to provide an accurate answer.

[0067] 102. Using a manifold learning enhancement algorithm, perform manifold embedding processing on the multimodal data in the personalized question understanding framework to reveal the intrinsic geometric structure of the multimodal data to optimize the information integration between the modalities in the multimodal data, use multi-task learning technology to construct a shared representation space, and simultaneously train multiple subtasks in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a cross-modal unified representation;

[0068] In this step, the manifold learning enhancement algorithm is a data analysis method that aims to reveal the low-dimensional intrinsic geometric structure in high-dimensional data sets. The algorithm retains the main features of the data through dimensionality reduction techniques while optimizing the information integration between different modalities.

[0069] Manifold embedding processing refers to the process of mapping high-dimensional data into a low-dimensional space so that the relative distances between data points remain as constant as possible. This method can help discover the intrinsic patterns of the data and improve the efficiency of data processing.

[0070] Intrinsic geometric structure refers to the spatial relationships and structural characteristics hidden behind multimodal data. Through manifold embedding processing, these characteristics can be displayed more clearly, thereby better understanding the essence of the data.

[0071] The shared representation space is a unified feature space in which data from different modalities are represented as vectors with similar features. Constructing such a space helps to process cross-modal data in a consistent manner.

[0072] The cross-modal unified representation is the final form of processed multimodal data in a common representation space. It integrates the data features of each modality and provides a unified data view to facilitate subsequent analysis and application.

[0073] In the embodiment of the present application, first, the system uses a manifold learning enhancement algorithm to perform manifold embedding processing on multimodal data in a personalized problem understanding framework; second, through this processing, the intrinsic geometric structure of the data is revealed, and the information integration between the modalities is optimized; third, multiple subtasks are trained simultaneously in the constructed shared representation space; finally, a unified cross-modal representation is generated as the basis for the next step of processing.

[0074] For example, continuing with the above example, assume that the system has generated a personalized question understanding framework that contains text and image features. The system first uses the manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the framework; secondly, through this processing, the intrinsic geometric structure of text and image data is revealed, and the information integration between the two is optimized; thirdly, multiple subtasks such as text classification and image recognition are trained simultaneously in the constructed shared representation space; finally, a cross-modal unified representation is generated to prepare for subsequent time series analysis and recursive processing.

[0075] 103. Using a recurrent neural network algorithm, recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, simulate the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, use a hierarchical clustering method to construct an organized information structure for the sequence information through a hierarchical aggregation method, and gradually merge the closest data points in the sequence information according to the similarity measurement in the organized information structure, provide deep knowledge mining, and generate an ordered information structure diagram;

[0076] In this step, the recurrent neural network algorithm is a neural network model specifically used to process sequence data. It is particularly suitable for capturing dependencies in time series. By introducing an internal state mechanism, it recursively updates the state of each time step and simulates the dynamic changes of time series.

[0077] Internal state refers to the memory saved by the recurrent neural network algorithm at each time step to pass on contextual information. The internal state allows the model to maintain memory of past information when processing long sequences, thereby better capturing time dependencies.

[0078] The time series feature set is a set of features extracted from time series data, which reflects the patterns and periodicity in the time series. They provide the basis for subsequent recursive processing.

[0079] Dynamic time changes refer to the characteristics of time series data evolving over time. The recurrent neural network algorithm can simulate this change and capture the complex dependencies between sequence information.

[0080] The organized information structure is a structured representation constructed by hierarchical aggregation of sequence information. It reflects the hierarchical relationship between data points. The hierarchical clustering method forms this structure by gradually merging the closest data points.

[0081] An ordered information structure diagram is a graphical representation of an organized information structure that shows the associations and hierarchical relationships between data points.

[0082] In the embodiment of the present application, first, the system applies time series feature extraction technology to generate a time series feature set; second, the recurrent neural network algorithm is used to recursively process the sequence information in the cross-modal unified representation and recursively update the internal state; third, the dynamic time changes of the sequence information are simulated to capture the dependency relationship; finally, the hierarchical clustering method is used to construct an organized information structure and generate an ordered information structure diagram.

[0083] For example, continuing with the previous example, assume that the system has generated a unified cross-modal representation. The system first applies the time series feature extraction technology to extract the time series feature set from the unified representation; secondly, it uses the recurrent neural network algorithm to recursively process these features, recursively update the internal state, and simulate the dynamic changes of the time series; thirdly, it captures the dependency between sequence information and generates a time series internal state diagram; finally, it uses the hierarchical clustering method to construct an organized information structure and generate an ordered information structure diagram to provide users with a detailed historical background analysis of tourist attractions.

[0084] 104. Introduce a context-aware feedback mechanism to monitor users' immediate reactions and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight assignment on the ordered information structure diagram, record users' immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

[0085] The context-aware feedback mechanism is a technology that monitors and responds to users' immediate reactions and environmental changes in real time. This mechanism can dynamically adjust the behavior of the system according to user behavior and environmental conditions to provide more personalized and accurate services.

[0086] Dynamic weight allocation refers to dynamically adjusting the weight values ​​of various factors based on their importance in different situations.

[0087] The intelligent data analysis and knowledge question-answering solution is a comprehensive solution generated based on a context-aware feedback mechanism. It not only takes into account the user's immediate feedback, but also combines historical data and current context to provide users with the most appropriate answers.

[0088] In the embodiment of the present application, first, the system introduces a context-aware feedback mechanism to monitor the user's immediate reaction and environmental changes in real time; secondly, dynamically adjust the answer strategy for the personalized question understanding framework; thirdly, evaluate and optimize the relevance and accuracy of the answer content by performing context-related dynamic weight allocation on the ordered information structure diagram; finally, record the user's immediate feedback on the answer content to generate intelligent data analysis and knowledge question and answer solutions.

[0089] For example, continuing with the above example, assuming that the system has generated an ordered information structure diagram. The system first introduces a context-aware feedback mechanism to monitor the user's immediate reaction and environmental changes in real time; secondly, dynamically adjust the answer strategy for the personalized question understanding framework to ensure that the answer is in line with the user's current interests and takes into account historical preferences; thirdly, through context-related dynamic weight allocation of the ordered information structure diagram, evaluate and optimize the relevance and accuracy of the answer content; finally, record the user's immediate feedback on the answer content, generate intelligent data analysis and knowledge question and answer solutions, and continuously improve service quality.

[0090] In summary, steps 101 to 104 cover the complete process from receiving a multimodal query request to generating intelligent data analysis and knowledge question-answering solutions, aiming to provide a multimodal large model-driven intelligent question-answering system to meet users' precise information needs in complex query scenarios.

[0091] In order to solve the complexity of multimodal data integration, in some embodiments, the use of the manifold learning enhancement algorithm in step 102 to perform manifold embedding processing on the multimodal data in the personalized question understanding framework includes: extracting the multimodal data characteristics in the personalized question understanding framework to analyze the potential associations between different data characteristics in the multimodal data characteristics and generate a multimodal association map; using the manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework into the multimodal association map to reveal the intrinsic geometric structure of the multimodal data to optimize the multimodal data. The information between the modalities is integrated to generate a manifold embedding representation; a shared representation space is constructed using multi-task learning technology, and multiple interrelated subtasks are designed for the manifold embedding representation. The interrelated subtasks are simultaneously trained in the shared representation space to improve the transmission efficiency of the personalized problem understanding framework and generate a multi-task collaborative model; the output results of each subtask in the multi-task collaborative model are obtained, mapped back to the shared representation space, and an adaptive regularization mechanism is introduced to dynamically perform regularization processing according to the output results of each subtask, so as to improve the generalization ability of the multi-task collaborative model and generate a cross-modal unified representation.

[0092] In this embodiment, the multimodal data characteristics refer to various data features extracted from the personalized question understanding framework, including sentiment analysis of text, color distribution and texture of images.

[0093] A multimodal association map is a graphical representation that shows the relationships between multimodal data features. By constructing this map, we can better capture the intrinsic connections between different modal data and provide a basis for subsequent data processing.

[0094] Manifold embedding representation is a data representation after being processed by the manifold learning enhancement algorithm. It retains the main features of the original data and reveals the intrinsic geometric structure of the data. This method helps to improve the efficiency and accuracy of data processing.

[0095] The multi-task collaborative model is the result of simultaneously training multiple subtasks in a shared representation space, where each subtask focuses on solving a specific type of problem, and all subtasks work together to improve the overall performance of the system.

[0096] The adaptive regularization mechanism is a technology that dynamically adjusts model parameters to prevent overfitting and improve the generalization ability of the model. It ensures that the model can maintain good performance in different situations by performing dynamic regularization based on the output results of the subtask.

[0097] The cross-modal unified representation is a comprehensive data representation that is ultimately generated, which integrates data features from different modalities and represents these features in a common space.

[0098] In an embodiment of the present application, first, the system extracts the multimodal data characteristics in the personalized problem understanding framework, analyzes the potential correlations between these characteristics, and generates a multimodal correlation map; secondly, the manifold learning enhancement algorithm is used to perform manifold embedding processing on the multimodal data, and the data is mapped to the multimodal correlation map, revealing its intrinsic geometric structure, optimizing the information integration between the modalities, and generating a manifold embedding representation; thirdly, multi-task learning technology is used to construct a shared representation space, and multiple interrelated subtasks are designed. These subtasks are trained simultaneously in the shared representation space to generate a multi-task collaborative model; finally, the output results of each subtask in the multi-task collaborative model are obtained, mapped back to the shared representation space, and an adaptive regularization mechanism is introduced. Regularization processing is dynamically performed according to the output results of each subtask, so as to improve the generalization ability of the multi-task collaborative model and generate a final cross-modal unified representation.

[0099] Here is a specific example:

[0100] For example, suppose a user uploads an artwork authentication request containing text descriptions and pictures, asking about the historical background and authenticity of a certain artwork. First, the system extracts multimodal data features in the personalized question understanding framework, such as descriptions in text and detailed features in pictures, analyzes the potential correlations between these features, and generates a multimodal correlation map; second, the manifold learning enhancement algorithm is used to perform manifold embedding processing on multimodal data, map the data to the multimodal correlation map, reveal its inherent geometric structure, optimize the integration between text and image data, and generate a manifold embedding representation; third, multi-task learning technology is used to construct a shared representation space, design multiple interrelated subtasks, such as artwork history analysis and authenticity identification, and train these subtasks simultaneously in the shared representation space to generate a multi-task collaborative model; finally, the output results of each subtask in the multi-task collaborative model are obtained and mapped back to the shared representation space, and an adaptive regularization mechanism is introduced to dynamically perform regularization processing according to the output results of each subtask, so as to improve the generalization ability of the multi-task collaborative model, generate the final cross-modal unified representation, and provide users with accurate artwork authentication services.

[0101] In order to solve the complexity of multimodal data integration, in some embodiments, the use of the manifold learning enhancement algorithm in step 102 to perform manifold embedding processing on the multimodal data in the personalized question understanding framework into the multimodal association map includes: introducing a similarity matrix to evaluate the multimodal correlation in the multimodal association map, identifying highly consistent parts in the multimodal association map, using local linear embedding technology to retain the local neighborhood structure during the identification process, and generating a consistent local neighborhood structure; using the manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework into the multimodal association map, and converting the consistent local neighborhood structure into a local linear embedding technology. The domain structure is used as the structural basis to reveal the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding intermediate result; for the manifold embedding intermediate result, a global information fusion mechanism is introduced, and the manifold embedding intermediate result is further optimized by combining the global context information in the personalized problem understanding framework to generate a global optimized embedding graph; a weighted average strategy is introduced to assign different weights according to the importance of different data points in the global optimized embedding graph, and a dimension selection technology is applied to select the dimension combination that is most representative of the multimodal data characteristics from the candidate dimensions generated by the global optimized embedding graph to generate a manifold embedding representation.

[0102] In this embodiment, the similarity matrix is ​​a tool for evaluating the correlation between different data points in a multimodal association graph.

[0103] A consistent local neighborhood structure refers to a data structure that retains local neighborhood relationships in a multimodal association graph. Through local linear embedding technology, it reveals the intrinsic geometric properties of the data while maintaining local features, ensuring that the relative position between each data point and its nearest neighbor remains unchanged.

[0104] The manifold embedding intermediate result is an intermediate representation generated after the preliminary manifold embedding processing. It reveals the intrinsic geometric structure of multimodal data based on the consistent local neighborhood structure and optimizes the information integration between the modalities.

[0105] The global information fusion mechanism is a technique that combines global contextual information in a personalized question understanding framework to further optimize the manifold embedding intermediate results. This approach ensures that the final generated embedding graph not only retains local features but also reflects the global structure.

[0106] The weighted average strategy is a method that assigns different weights according to the importance of different data points in the global optimization embedding graph. By applying dimension selection technology, the most representative dimension combination is selected from the candidate dimensions to generate the final manifold embedding representation.

[0107] In an embodiment of the present application, first, a similarity matrix is ​​introduced to evaluate the multimodal correlation in the multimodal association graph, and the highly consistent parts in the graph are identified; secondly, a local linear embedding technique is used to retain the local neighborhood structure during the recognition process and generate a consistent local neighborhood structure; thirdly, a manifold learning enhancement algorithm is used to perform manifold embedding processing on the multimodal data into the multimodal association graph, and based on the consistent local neighborhood structure, the intrinsic geometric structure of the data is revealed to generate a manifold embedding intermediate result; finally, a global information fusion mechanism is introduced for the manifold embedding intermediate result, and the global context information in the personalized problem understanding framework is combined to further optimize the intermediate result and generate a globally optimized embedding graph, and a weighted average strategy is introduced to assign different weights according to the importance of different data points, and a dimension selection technique is applied to select the most representative dimension combination to generate the final manifold embedding representation.

[0108] Here is a specific example:

[0109] For example, suppose a user uploads an educational material query request containing text descriptions and video clips, asking for a detailed explanation and application scenario of a certain subject knowledge point. First, the system introduces a similarity matrix to evaluate the multimodal correlation in the multimodal association graph and identify the highly consistent parts of the graph; secondly, the local linear embedding technology is used to retain the local neighborhood structure during the recognition process and generate a consistent local neighborhood structure; thirdly, the manifold learning enhancement algorithm is used to perform manifold embedding processing on the multimodal data into the multimodal association graph, based on the consistent local neighborhood structure, revealing the intrinsic geometric structure of the data and generating a manifold embedding intermediate result; finally, for the manifold embedding intermediate result, a global information fusion mechanism is introduced, combined with the global context information in the personalized problem understanding framework, the intermediate result is further optimized, a globally optimized embedding graph is generated, and a weighted average strategy is introduced to assign different weights according to the importance of different data points, and the dimension selection technology is applied to select the most representative dimension combination to generate the final manifold embedding representation, providing users with accurate educational material query services.

[0110] In order to solve the problem of multimodal data integration and transmission efficiency, in some embodiments, the multi-task learning technology is used in step 102 to construct a shared representation space, and multiple interrelated subtasks are designed for the manifold embedding representation. The interrelated subtasks are trained simultaneously in the shared representation space to improve the transmission efficiency of the personalized problem understanding framework, and a multi-task collaborative model is generated, including: identifying natural clusters in the manifold embedding representation by a clustering method to obtain concentrated areas in the manifold embedding representation, evaluating the internal consistency and external separability of each cluster in the natural cluster, and generating a manifold embedding cluster analysis result; using multi-task learning technology to construct a shared representation space based on the manifold embedding cluster analysis result, for The manifold embedding representation designs multiple interrelated subtasks, and the interrelated subtasks are simultaneously trained in the shared representation space to improve the transfer efficiency of the personalized problem understanding framework and generate a multi-task learning architecture; for the multi-task learning architecture, a task correlation matrix is ​​introduced to quantify the task correlation in the multi-task learning architecture to enhance the synergy between the subtasks in the multi-task learning architecture and generate a task collaborative optimization graph; in the shared representation space, an initial weight is assigned to each collaborative task in the task collaborative optimization graph, the initial weight is dynamically adjusted through a back-propagation algorithm, and an early stopping mechanism is introduced to prevent overfitting during the dynamic adjustment process and generate a multi-task collaborative model.

[0111] In this embodiment, natural clusters refer to concentrated areas of data points identified from the manifold embedding representation by clustering methods. These clusters reflect natural groupings in the intrinsic structure of the data and help to evaluate the internal consistency and external separability between different areas.

[0112] The manifold embedding cluster analysis result is generated by evaluating natural clusters. It shows the internal consistency and external separation of each cluster, helping the system to better understand the data distribution and characteristics.

[0113] The shared representation space is a unified feature space in which data from different modalities are represented as vectors with similar features.

[0114] The multi-task learning architecture is a learning framework based on the results of manifold embedding cluster analysis, which contains multiple interrelated subtasks. Each subtask focuses on solving a specific type of problem, and all subtasks work together to improve the overall performance of the system.

[0115] The task correlation matrix is ​​a quantitative tool used to measure the correlation between subtasks in a multi-task learning architecture. By introducing this matrix, the synergy between subtasks can be enhanced and the learning effect of the model can be improved.

[0116] The task co-optimization graph is a graphical representation generated from the task dependency matrix, showing the collaborative relationship between subtasks. This method helps to further optimize the multi-task learning process and ensure that the subtasks can work together effectively.

[0117] The multi-task collaborative model is the final model after training. It combines the advantages of multiple sub-tasks and improves the system's delivery efficiency and accuracy when processing complex multimodal queries.

[0118] In an embodiment of the present application, first, a clustering method is used to identify natural clusters in the manifold embedding representation, concentrated areas are obtained, and the internal consistency and external separability of each cluster are evaluated to generate a manifold embedding cluster analysis result; secondly, a multi-task learning technique is used to construct a shared representation space based on the manifold embedding cluster analysis result, and multiple interrelated subtasks are designed. These subtasks are trained simultaneously in the shared representation space to generate a multi-task learning architecture; thirdly, for the multi-task learning architecture, a task correlation matrix is ​​introduced to quantify the relationship between tasks, enhance the synergy between subtasks, and generate a task collaborative optimization graph; finally, in the shared representation space, initial weights are assigned to each collaborative task in the task collaborative optimization graph, these weights are dynamically adjusted through the back-propagation algorithm, and an early stopping mechanism is introduced to prevent overfitting, thereby generating a multi-task collaborative model.

[0119] Here is a specific example:

[0120] For example, suppose a user uploads a music review request containing a text description and an audio clip, asking about the emotional expression and playing skills of a certain song. First, the system uses a clustering method to identify natural clusters in the manifold embedding representation, obtains concentrated areas, and evaluates the internal consistency and external separability of each cluster to generate the manifold embedding cluster analysis results; secondly, the multi-task learning technology is used to build a shared representation space based on the manifold embedding cluster analysis results, and multiple interrelated subtasks such as sentiment analysis and playing skill recognition are designed. These subtasks are trained simultaneously in the shared representation space to generate a multi-task learning architecture; thirdly, for the multi-task learning architecture, a task correlation matrix is ​​introduced to quantify the relationship between tasks, enhance the synergy between subtasks, and generate a task synergy optimization graph; finally, in the shared representation space, initial weights are assigned to each collaborative task in the task synergy optimization graph, and these weights are dynamically adjusted through the back-propagation algorithm, and an early stopping mechanism is introduced to prevent overfitting, so as to generate a multi-task collaborative model and provide users with accurate music review services.

[0121] In order to solve the problem of capturing the temporal dependency and dependency relationship of sequence information, in some embodiments, the use of a recurrent neural network algorithm in step 103 recursively processes the sequence information in the cross-modal unified representation, including: applying a temporal feature extraction technique to extract the temporal dependency component in the cross-modal unified representation, analyzing the temporal pattern and periodic features in the temporal dependency component, and generating a time series feature set; using a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation, recursively updating the internal state of the sequence information, inputting the internal state into the time series feature set, and simulating the sequence information. The dynamic time changes of the sequence information are captured to capture the dependency relationship between the sequence information and generate the internal state diagram of the time series; a hierarchical clustering method is used to construct an organized information structure for the sequence information and the internal state diagram of the time series through hierarchical aggregation; according to the similarity measurement in the organized information structure, the closest data points in the sequence information are gradually merged to provide deep knowledge mining and generate a hierarchical clustering structure diagram; the key nodes and paths in the hierarchical clustering structure diagram are identified, and a graph traversal algorithm is applied to explore layer by layer from the root node to the leaf node in the key node, and the important transmission paths of the ordered information in the layer-by-layer exploration process are recorded to generate an ordered information structure diagram.

[0122] In this embodiment, the time series feature extraction technology is a technology for extracting time-dependent components from data.

[0123] The time series feature set is a data set generated by time series feature extraction technology, which contains the time patterns and periodic features in the time-dependent components. These feature sets provide important input information for the recurrent neural network algorithm, helping it to better simulate dynamic time changes.

[0124] The time series internal state diagram is a graphical representation generated by the recurrent neural network algorithm, which shows how the internal state of the sequence information changes over time. This diagram helps capture the dependencies between sequence information and provides visualization tools to assist analysis.

[0125] The hierarchical clustering structure diagram is an organized information structure diagram constructed by hierarchical aggregation, which shows the merging process between the closest data points in the sequence information. This method can provide in-depth knowledge mining and reveal the inherent hierarchical relationship of the data.

[0126] The ordered information structure diagram is the final result generated by exploring the key nodes and paths in the hierarchical clustering structure diagram layer by layer through a graph traversal algorithm. It records important transmission paths, shows the ordered information flow between data points, and provides users with clear analysis results.

[0127] In the embodiment of the present application, first, the time series feature extraction technology is applied to extract the time-dependent components from the cross-modal unified representation, analyze the time pattern and periodic features therein, and generate a time series feature set; secondly, the recurrent neural network algorithm is used to recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, input the internal state into the time series feature set, simulate the dynamic time change of the sequence information, and generate a time series internal state diagram; thirdly, the hierarchical clustering method is adopted to construct an organized information structure for the sequence information and the time series internal state diagram through hierarchical aggregation, gradually merge the closest data points according to the similarity measurement, provide deep knowledge mining, and generate a hierarchical clustering structure diagram; finally, the key nodes and paths in the hierarchical clustering structure diagram are identified, and the graph traversal algorithm is applied to explore layer by layer from the root node to the leaf node, record the important transmission paths of the ordered information, and generate an ordered information structure diagram.

[0128] Here is a specific example:

[0129] For example, suppose a user uploads a fitness training request containing text descriptions and video clips, asking about the standard posture and common mistakes of a certain fitness action. First, the system applies temporal feature extraction technology to extract time-dependent components from the cross-modal unified representation, analyzes the temporal patterns and periodic features in these components, and generates a time series feature set; secondly, the system uses a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation, update the internal state of the sequence information, and input these internal states into the time series feature set to simulate its dynamic time changes to capture the dependency between sequence information and generate a time series internal state graph; thirdly, the hierarchical clustering method is used to construct an organized information structure for the sequence information and the time series internal state graph through hierarchical aggregation. According to the similarity measurement in the organized information structure, the closest data points are gradually merged to provide deep knowledge mining and generate a hierarchical clustering structure graph; finally, the key nodes and paths in the hierarchical clustering structure graph are identified, and the graph traversal algorithm is applied to explore layer by layer from the root node to the leaf node in the key node, record the important transmission path of the ordered information in the layer-by-layer exploration process, generate an ordered information structure graph, and provide users with detailed fitness training guidance.

[0130] In order to solve the problem of capturing complex time dependencies and abnormal patterns in sequence information, in some embodiments, the use of a recurrent neural network algorithm in combination with a sliding window mechanism and a time-frequency transformation technology in step 103 includes: applying a sliding window mechanism to extract local time series features in the time series feature set, combining the time-frequency transformation technology to capture global periodic patterns based on the local time series features, and generating a refined time series feature set; using a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation and the refined time series feature set, recursively updating the internal state of the sequence information, inputting the internal state into the time series feature set, simulating the dynamic time changes of the sequence information to capture the dependency relationship between the sequence information, and generating a recursively updated state record; mapping the internal state of each time step in the recursively updated state record to the corresponding time point, and applying an anomaly detection algorithm to identify and mark potential abnormal mapping conversions to generate a time series state mapping table; integrating the internal state information of all time steps in the time series state mapping table, applying visualization technology to convert the mapping results in the time series state mapping table into an intuitive graphical representation, and generating a time series internal state diagram.

[0131] In this embodiment, the sliding window mechanism is a method for extracting local time series features from a time series feature set.

[0132] Time-frequency transformation technology refers to the technology of converting time domain signals into frequency domain, which can help identify periodic components in time series and provide information about the global pattern of time series.

[0133] The refined time series feature set is an enhanced feature set generated by the sliding window mechanism and time-frequency transformation technology. It not only contains local time series features, but also captures global periodic patterns, allowing the model to better understand the dynamic changes in sequence information.

[0134] The recursive update state record is an internal state record generated after recursively updating the sequence information in the cross-modal unified representation and the refined time series feature set during the recursive processing. This record tracks the changes at each time step and helps capture the dependencies between sequences.

[0135] The time series state map is a table created by mapping the internal state of each time step in the recursively updated state record to the corresponding time point. An anomaly detection algorithm is applied to mark any possible abnormal transitions, which helps to identify potential abnormal situations or unusual behaviors.

[0136] The time series internal state diagram is the product of converting the mapping results in the time series state mapping table into intuitive graphical representations through visualization technology. These graphics provide a visual display of the changes in the internal state of the time series over time, which is convenient for analysts to understand and interpret.

[0137] In an embodiment of the present application, first, a sliding window mechanism is applied to extract local time series features from a time series feature set, and combined with time-frequency transformation technology, global periodic patterns are captured based on local time series features to generate a refined time series feature set; secondly, a recurrent neural network algorithm is used to recursively process the sequence information in the cross-modal unified representation and the refined time series feature set, recursively update the internal state of the sequence information, input the internal state to the time series feature set, simulate the dynamic time changes of the sequence information, so as to capture the dependency relationship between the sequence information, and generate a recursive update state record; thirdly, the internal state of each time step in the recursively updated state record is mapped to the corresponding time point, and an anomaly detection algorithm is applied to identify and mark potential abnormal mapping conversions to generate a time series state mapping table; finally, the internal state information of all time steps in the time series state mapping table is integrated, and visualization technology is applied to convert the mapping results into an intuitive graphical representation to generate a time series internal state diagram.

[0138] Here is a specific example:

[0139] Suppose that in an intelligent medical monitoring system, the user wants to analyze the patient's electrocardiogram data to identify heart problems such as arrhythmias. First, the system applies a sliding window mechanism to extract local time series features from the time series feature set of the electrocardiogram data, and combines the time-frequency transformation technology to capture the global periodic pattern and generate a refined time series feature set; secondly, the recurrent neural network algorithm is used to recursively process the sequence information of the electrocardiogram data and the refined time series feature set, recursively update the internal state of the sequence information, simulate its dynamic time changes, capture the dependency between the sequence information, and generate a recursive update state record; thirdly, the internal state of each time step in the recursive update state record is mapped to the corresponding time point, and the anomaly detection algorithm is applied to identify and mark potential abnormal mapping transformations to generate a time series state mapping table; finally, the internal state information of all time steps in the time series state mapping table is integrated, and the visualization technology is applied to convert the mapping results into an intuitive graphical representation to generate a time series internal state diagram to help doctors diagnose heart diseases more accurately.

[0140] In order to solve the problem of the influence of user's immediate reaction and environmental changes on the question-answering strategy, in some embodiments, the introduction of the context-aware feedback mechanism in step 104 monitors the user's immediate reaction and environmental changes in real time, and dynamically adjusts the answer strategy for the personalized question understanding framework, including: parsing the relationship between the nodes and the edges in the ordered information structure graph to identify the logical connection between different information elements in the ordered information structure graph, extracting the information units and interaction modes represented by each node in the ordered information structure graph through pattern recognition technology, and generating a node and edge information table; introducing a context-aware feedback mechanism to capture the user's immediate reaction and the surrounding environmental changes, and converting the immediate reaction and the surrounding environmental changes into a computable feature vector in combination with the node and edge information table to generate a computable feature vector group; dynamically adjusting the answer strategy for the personalized question understanding framework, by performing context-related dynamic weight allocation on the ordered information structure graph, dynamically weighting the computable feature vector group, and generating a dynamic weighted feature vector; evaluating the relevance and accuracy of the answer content in the answer strategy, optimizing the evaluation result in combination with the dynamic weighted feature vector, recording the user's immediate feedback on the answer content, and generating an intelligent data analysis and knowledge question-answering solution.

[0141] In this embodiment, the ordered information structure diagram is a graphical representation that shows the correlation and hierarchical relationship between data points. It is composed of nodes (representing information units) and edges (representing logical connections between information units), providing a deep knowledge mining tool.

[0142] The node and edge information table is a table generated by parsing the relationship between nodes and edges in an ordered information structure graph. This table records in detail the information units represented by each node and how they interact with each other, helping the system to better understand the logical connections between information.

[0143] Pattern recognition technology is a technology used to extract meaningful patterns from complex data. By applying this technology, we can identify the information units represented by each node in the ordered information structure diagram and describe the way they interact with each other.

[0144] The context-aware feedback mechanism is a technology that monitors users' immediate reactions and changes in the surrounding environment in real time. It can capture users' immediate feedback and dynamic changes in the environment, and convert this information into feature vectors that can be used for calculation.

[0145] The computable feature vector group is a group of feature vectors generated by the context-aware feedback mechanism, which reflects the user's immediate reaction and environmental changes. These vectors can be used to dynamically adjust the answer strategy in subsequent processing.

[0146] The dynamically weighted feature vector is the result of context-dependent dynamic weight assignment to a group of computable feature vectors. In this way, the system can adjust the importance of each feature according to the current situation, thereby optimizing the answer strategy.

[0147] The intelligent data analysis and knowledge question-answering solution is a comprehensive solution that combines users' immediate feedback, historical data and current situations to provide users with the most appropriate answers.

[0148] In an embodiment of the present application, first, the relationship between nodes and edges in an ordered information structure graph is parsed, the logical connection between different information elements is identified, and the information units and interaction modes represented by each node are extracted through pattern recognition technology to generate a node and edge information table; secondly, a context-aware feedback mechanism is introduced to capture the user's immediate reaction and changes in the surrounding environment, and these reactions and changes are converted into computable feature vectors in combination with the node and edge information tables to generate a computable feature vector group; thirdly, the answer strategy for the personalized question understanding framework is dynamically adjusted, and the computable feature vector group is dynamically weighted by context-related dynamic weight allocation to the ordered information structure graph to generate a dynamic weighted feature vector; finally, the relevance and accuracy of the answer content in the answer strategy are evaluated, the evaluation results are optimized in combination with the dynamic weighted feature vector, the user's immediate feedback on the answer content is recorded, and an intelligent data analysis and knowledge question and answer solution is generated.

[0149] Here is a specific example:

[0150] For example, in an intelligent traffic navigation system, users want to obtain the best route suggestions to avoid traffic congestion. First, the system analyzes the relationship between nodes and edges in the ordered information structure graph, identifies the logical connection between different information elements, extracts the information units and interaction modes represented by each node through pattern recognition technology, and generates node and edge information tables; secondly, the context-aware feedback mechanism is introduced to capture the user's immediate reaction (such as the user's selection or cancellation of a recommended route) and changes in the surrounding environment (such as real-time traffic status updates), and these reactions and changes are converted into computable feature vectors in combination with the node and edge information tables to generate a computable feature vector group; thirdly, the answer strategy for the personalized question understanding framework is dynamically adjusted, and the computable feature vector group is dynamically weighted by context-related dynamic weight allocation to the ordered information structure graph to generate a dynamic weighted feature vector; finally, the relevance and accuracy of the answer content in the answer strategy are evaluated, and the evaluation results are optimized in combination with the dynamic weighted feature vector, and the user's immediate feedback on the answer content is recorded, and an intelligent data analysis and knowledge question and answer solution is generated to ensure that the route suggestions provided are in line with the current traffic conditions and meet the user's preferences.

[0151] The present application considers that in order to solve the problems in the prior art of low efficiency in integrating multimodal data and difficulty in capturing the intrinsic geometric structure, the embodiment of the invention proposes this optional solution to solve the technical problem of improving the accuracy and generalization ability of multimodal data analysis, and thus proposes a new optional solution, which includes:

[0152] Using the manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is processed into the multimodal association graph by manifold embedding, and the consistent local neighborhood structure is used as the structural basis to reveal the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding intermediate result, including:

[0153] By principal component analysis, high-dimensional multimodal data in the consistent local neighborhood structure is mapped to a low-dimensional space, and a k-nearest neighbor graph is constructed to obtain similarity weights in the low-dimensional space to generate a local neighborhood structure consistency loss function;

[0154] The local neighborhood structure consistency loss function is calculated using the following formula:

[0155] ;

[0156] in, is the local neighborhood structure consistency loss function; is the number of multimodal data samples; For sample of Neighbor set; For sample and The similarity weight between them is used to measure the correlation between samples in the local neighborhood; For sample and The square of the Euclidean distance between them is used to maintain the consistency of the local neighborhood structure; is the control coefficient of the logarithmic term; is the nonlinear adjustment coefficient, which is used to adjust the speed of exponential decay;

[0157] Applying exponential decay functions and periodic functions to enhance the ability of the local neighborhood structure consistency loss function to capture complex geometric structures, adding regularization terms to prevent overfitting, and improving the generalization performance of the local neighborhood structure consistency loss function to generate a manifold embedding energy function;

[0158] The manifold embedding energy function is calculated using the following formula:

[0159] ;

[0160] in, Embed energy function for manifold; is the exponential decay coefficient, which is used to adjust the strength of the nonlinear relationship; is the exponential decay rate coefficient; is the sine wave amplitude coefficient, used to introduce periodic changes; is the sine wave period adjustment coefficient; Is a balance parameter used to adjust the similarity weight The impact of For sample and The similarity weight between them; is a regularization term used to prevent overfitting; is the regularization coefficient, which is used to control the influence of the regularization term on the energy function; For low-dimensional representation and The square of the Euclidean distance between them is used to maintain the consistency of the internal geometric structure; and They are respectively and A low-dimensional representation of samples;

[0161] The gradient descent method is used to iteratively update the positions of sample points in the manifold embedding energy function in the low-dimensional space to converge to the optimal solution, and the energy change after each iteration during the iterative update process is monitored to generate an intermediate result of manifold embedding.

[0162] This method aims to capture the intrinsic geometric structure of multimodal data and optimize the information integration between modes by manifold embedding. Based on principal component analysis and k-nearest neighbor graph, a local neighborhood structure consistency loss function is constructed, and exponential decay function and periodic function are further introduced to enhance the ability to capture complex geometric structures. At the same time, regularization terms are added to prevent overfitting. Finally, the gradient descent method is used to iteratively update the position of sample points in low-dimensional space to ensure that the energy function converges to the optimal solution and generates manifold embedding intermediate results. This method not only improves the accuracy of multimodal data analysis, but also enhances the generalization ability and response speed of the system, providing a solid foundation for intelligent data analysis and knowledge question answering.

[0163] Suppose in an intelligent customer support system based on a multimodal large model, the system needs to process customers' multimedia queries in real time to provide accurate answers to questions and personalized service recommendations;

[0164] The customer service interaction data within a specific time period is selected as the optimized time series feature set, which contains 300 samples (i.e. ), each sample has 80 feature dimensions, k is 5 (that is, each sample considers its 5 nearest neighbors); assuming that the average similarity weight calculated by cosine similarity is 0.85; spatial attenuation coefficient , nonlinear adjustment coefficient ;

[0165] ;

[0166] Assuming an exponential decay coefficient ; Sine wave amplitude coefficient , sine wave period adjustment coefficient ; Balance parameters ; Regularization coefficient ; Assume that for the above sample, find 200 time series feature sets that match it (i.e. ), and the regularization term Select L2 regularization;

[0167] ;

[0168] Assuming that the threshold is set to 0.85, since the calculated result 0.89 is greater than the set threshold, it shows that the intelligent customer support system has high effectiveness and reliability. This is because the higher energy function value reflects that the model can effectively capture the intrinsic geometric structure of multimodal customer service interaction data under the current conditions, while ensuring the quality of low-dimensional representation. This shows that the model can better retain the characteristics of the original data, which helps to more accurately understand the customer's problem pattern and service needs, thereby providing personalized solutions. Through the above steps, the effective dimensionality reduction and retention of the intrinsic geometric structure of the customer service interaction data are ensured, and the accuracy and response speed of the intelligent customer support system are improved.

[0169] This application considers that in order to solve the problem of insufficient capture of time dependency and complex sequence information of multimodal data in the prior art, the embodiment of the invention proposes this optional solution to solve the technical problem of improving the accuracy of time series analysis and the generalization ability of the model, and thus proposes a new optional solution, which includes:

[0170] Using a recurrent neural network algorithm, the sequence information in the cross-modal unified representation is combined with the refined time series feature set for recursive processing, the internal state of the sequence information is recursively updated, the internal state is input into the time series feature set, the dynamic time change of the sequence information is simulated to capture the dependency relationship between the sequence information, and a recursive update state record is generated, including:

[0171] Based on, by introducing multi-scale feature extraction technology, the feature changes at different time scales in the refined time series feature set are captured, and a local consistency preservation graph is constructed to calculate the similarity matrix of the feature changes to generate an internal state;

[0172] The internal state is calculated using the following formula:

[0173] ;

[0174] in, is the time step The internal state of the time series information; is the time step The internal state of the time series information; is the activation function; is the input weight matrix, connecting the input features of the current time step and internal state; is the recursive weight matrix, connecting the internal state of the previous time step and the internal state of the current time step; is the bias vector, used to adjust the initial value of the internal state; is the logarithmic adjustment coefficient; is the distance attenuation coefficient; The square of the Euclidean distance between the input features of the current time step and the internal state of the previous time step is used to maintain local consistency; is the refined time series feature set of the sequence information in the cross-modal unified representation at time step t;

[0175] Based on the internal state, the retention ratio of the memory information of the previous time step is controlled by a forget gate, the new information of the current time step is selected by an input gate, an exponential decay term is used to capture the distance-based change relationship, and a hyperbolic tangent activation function is introduced for nonlinear transformation to generate a memory cell state;

[0176] The memory cell state is calculated using the following formula:

[0177] ;

[0178] in, is the time step The state of memory cells; is the activation function; and They are the weight matrices of the forget gate and the input gate respectively; and are the bias vectors of the forget gate and input gate respectively; is the element-wise multiplication operator; tanh is the hyperbolic tangent activation function; is the weight matrix of the candidate memory cell state; is the bias vector of the candidate memory cell state; is the index adjustment coefficient; is the distance attenuation coefficient; The squared Euclidean distance between the input features at the current time step and the internal state at the current time step; is the refined time series feature set of the sequence information in the cross-modal unified representation at time step t; is the time step The internal state of the time series information; is the time step The state of memory cells;

[0179] The time-dependency analysis is applied to evaluate the correlation strength of the memory cell state at each time step in the memory cell state, and a recursive update formula is introduced to dynamically adjust the update mode of the internal state according to the memory cell state and the current input feature to generate a recursive update state record.

[0180] This method aims to capture the complex relationship between time dependencies and sequence information by recursively processing the time series features of multimodal data, and optimize the information integration between the modalities. Based on the recurrent neural network algorithm, the internal state of the sequence information is recursively updated by combining the sequence information in the cross-modal unified representation with the refined time series feature set. This method not only improves the accuracy of time series analysis, but also enhances the generalization ability and response speed of the system, providing a solid foundation for intelligent data analysis and knowledge question answering.

[0181] Suppose in an intelligent academic literature retrieval and question-answering system based on a multimodal large model, the system needs to process multimedia queries uploaded by users in real time to provide accurate knowledge answers and data insights;

[0182] Assume that the input weight matrix and the recursive weight matrix Obtained through training, the bias vector Initially set to zero, logarithmic adjustment coefficient , distance attenuation coefficient ;

[0183] ;

[0184] Assume that the weight matrices of the forget gate and the input gate are Obtained through training, the bias vector Initially set to zero, the weight matrix of the candidate memory cell state Obtained through training, the bias vector Initially set to zero, exponential adjustment coefficient , distance attenuation coefficient ;

[0185] ;

[0186] Assuming that the threshold is set to 0.9, since the calculated result 0.93 is greater than the set threshold, it shows that the intelligent academic literature retrieval and question-answering system has high effectiveness and reliability. This is because the higher memory cell state value reflects that the model can effectively capture the time dependency and sequence information of multimodal query data under the current conditions, while ensuring the quality of the internal state. This shows that the model can better understand the user's query intention and data pattern, which helps to more accurately match relevant academic literature and provide in-depth data analysis. Through the above steps, the effective processing of user query data and the retention of the intrinsic geometric structure are ensured, which improves the accuracy and response speed of the system. This enables researchers to obtain the latest research results they need more quickly, improve research efficiency, and enhance the intelligence and service level of the entire academic literature retrieval system.

[0187] Figure 2A structural diagram of an intelligent data analysis and knowledge question answering system based on a multimodal large model is provided for the embodiment of the present application, such as Figure 2 As shown, the device comprises:

[0188] The receiving module 21 is used to receive and parse the multimodal query request from the user, perform semantic parsing and intent recognition technology on the multimodal query request, determine the core of the question and the required knowledge field in the multimodal query request, and generate a personalized question understanding framework; the multimodal query request includes a combination of data in multiple forms such as text, image, audio and video;

[0189] A processing module 22 is used to use a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework, reveal the intrinsic geometric structure of the multimodal data, optimize the information integration between the modes in the multimodal data, use multi-task learning technology to construct a shared representation space, and simultaneously train multiple subtasks in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a cross-modal unified representation;

[0190] A simulation module 23 is used to use a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, simulate the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, use a hierarchical clustering method to construct an organized information structure for the sequence information through a hierarchical aggregation method, and gradually merge the closest data points in the sequence information according to the similarity measurement in the organized information structure, provide deep knowledge mining, and generate an ordered information structure diagram;

[0191] The monitoring module 24 is used to introduce a context-aware feedback mechanism, monitor the user's immediate reaction and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

[0192] Figure 2 The intelligent data analysis and knowledge question answering system based on a multimodal large model can be executed Figure 1 The implementation principle and technical effect of the intelligent data analysis and knowledge question answering method based on a multimodal large model described in the illustrated embodiment will not be repeated. The specific manner in which each module and unit performs operations in the intelligent data analysis and knowledge question answering system based on a multimodal large model in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0193] In one possible design, Figure 2 The intelligent data analysis and knowledge question answering system based on a multimodal large model of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0194] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0195] The processing component 32 is used to: receive and parse a multimodal query request from a user, perform semantic parsing and intent recognition technology on the multimodal query request, determine the problem core and required knowledge domain in the multimodal query request, and generate a personalized problem understanding framework; the multimodal query request includes a combination of data in multiple forms such as text, image, audio and video; use a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized problem understanding framework to reveal the inherent geometric structure of the multimodal data to optimize the information integration between the modes in the multimodal data, use multi-task learning technology to construct a shared representation space, and simultaneously train multiple subtasks in the shared representation space to improve the transmission efficiency of the personalized problem understanding framework and generate a cross-modal unified representation; use a recurrent neural network algorithm to embed the cross-modal unified representation The sequence information in the display is recursively processed, the internal state of the sequence information is recursively updated, the dynamic time change of the sequence information is simulated to capture the dependency relationship between the sequence information, and the hierarchical clustering method is adopted to construct an organized information structure for the sequence information through hierarchical aggregation. According to the similarity measurement in the organized information structure, the closest data points in the sequence information are gradually merged to provide deep knowledge mining and generate an ordered information structure diagram; a context-aware feedback mechanism is introduced to monitor the user's immediate reaction and environmental changes in real time, and the answer strategy for the personalized question understanding framework is dynamically adjusted. By performing context-related dynamic weight allocation on the ordered information structure diagram, the relevance and accuracy of the answer content in the answer strategy are evaluated and optimized, the user's immediate feedback on the answer content is recorded, and an intelligent data analysis and knowledge question and answer solution is generated.

[0196] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0197] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0198] Of course, the computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0199] The input / output interface provides an interface between the processing component and the peripheral interface module, which may be an output device, an input device, etc.

[0200] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0201] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0202] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment is an intelligent data analysis and knowledge question-answering method based on a multimodal large model.

[0203] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0204] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0205] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An intelligent data analysis and knowledge question answering method based on a multimodal large model, characterized in that: include: Receive and parse multimodal query requests from users, perform semantic parsing and intent recognition technology on the multimodal query requests, determine the core of the question and the required knowledge domain in the multimodal query requests, and generate a personalized question understanding framework; the multimodal query requests contain a combination of data in multiple forms such as text, images, audio and video; Using a manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is subjected to manifold embedding processing to reveal the intrinsic geometric structure of the multimodal data so as to optimize the information integration between the modalities in the multimodal data. A shared representation space is constructed using multi-task learning technology, and multiple subtasks are trained simultaneously in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a unified cross-modal representation. The generating cross-modal unified representation process generates a manifold embedding representation, and the generating manifold embedding representation process generates a manifold embedding intermediate result; The generating manifold embedding intermediate result comprises: By principal component analysis, high-dimensional multimodal data in the consistent local neighborhood structure are mapped to a low-dimensional space, and a k-nearest neighbor graph is constructed to obtain similarity weights in the low-dimensional space to generate a local neighborhood structure consistency loss function; The local neighborhood structure consistency loss function is calculated using the following formula: ; in, is the local neighborhood structure consistency loss function; is the number of multimodal data samples; For sample of Neighbor set; For sample and The similarity weight between them is used to measure the correlation between samples in the local neighborhood; For sample and The square of the Euclidean distance between them is used to maintain the consistency of the local neighborhood structure; is the control coefficient of the logarithmic term; is the nonlinear adjustment coefficient, which is used to adjust the speed of exponential decay; Applying exponential decay functions and periodic functions to enhance the ability of the local neighborhood structure consistency loss function to capture complex geometric structures, adding regularization terms to prevent overfitting, and improving the generalization performance of the local neighborhood structure consistency loss function to generate a manifold embedding energy function; The manifold embedding energy function is calculated using the following formula: ; in, Embed energy function for manifold; is the exponential decay coefficient, which is used to adjust the strength of the nonlinear relationship; is the exponential decay rate coefficient; is the sine wave amplitude coefficient, used to introduce periodic changes; is the sine wave period adjustment coefficient; Is a balance parameter used to adjust the similarity weight The impact of For sample and The similarity weight between them; is a regularization term used to prevent overfitting; is the regularization coefficient, which is used to control the influence of the regularization term on the energy function; For low-dimensional representation and The square of the Euclidean distance between them is used to maintain the consistency of the internal geometric structure; and They are respectively and A low-dimensional representation of samples; Adopting the gradient descent method, iteratively updating the positions of the sample points in the manifold embedding energy function in the low-dimensional space to converge to the optimal solution, monitoring the energy change after each iteration during the iterative update process, and generating the manifold embedding intermediate result; Using a recurrent neural network algorithm, recursively processing the sequence information in the cross-modal unified representation, recursively updating the internal state of the sequence information, simulating the dynamic time changes of the sequence information to capture the dependencies between the sequence information, using a hierarchical clustering method, constructing an organized information structure for the sequence information through hierarchical aggregation, and gradually merging the closest data points in the sequence information according to the similarity measure in the organized information structure, providing deep knowledge mining, and generating an ordered information structure diagram; A context-aware feedback mechanism is introduced to monitor users' immediate reactions and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record users' immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

2. The method according to claim 1, characterized in that: The method uses a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework to reveal the intrinsic geometric structure of the multimodal data so as to optimize the information integration between the modes in the multimodal data, uses multi-task learning technology to construct a shared representation space, and simultaneously trains multiple subtasks in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a cross-modal unified representation, including: Extracting multimodal data features in the personalized question understanding framework to analyze potential associations between different data features in the multimodal data features and generate a multimodal association map; Using a manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is processed into the multimodal association graph by manifold embedding, revealing the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding representation; A shared representation space is constructed using multi-task learning technology, multiple interrelated subtasks are designed for the manifold embedding representation, and the interrelated subtasks are simultaneously trained in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a multi-task collaborative model; The output results of each subtask in the multi-task collaborative model are obtained, mapped back to the shared representation space, and an adaptive regularization mechanism is introduced to dynamically perform regularization processing according to the output results of each subtask, so as to improve the generalization ability of the multi-task collaborative model and generate a cross-modal unified representation.

3. The method according to claim 2, characterized in that The method uses a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework into the multimodal association graph, reveals the intrinsic geometric structure of the multimodal data, optimizes the information integration between the modes in the multimodal data, and generates a manifold embedding representation, including: A similarity matrix is ​​introduced to evaluate the multimodal correlation in the multimodal association map, and a highly consistent part in the multimodal association map is identified. A local linear embedding technique is used to retain the local neighborhood structure during the identification process to generate a consistent local neighborhood structure. Using a manifold learning enhancement algorithm, the multimodal data in the personalized question understanding framework is subjected to manifold embedding processing into the multimodal association graph, and the consistent local neighborhood structure is used as a structural basis to reveal the intrinsic geometric structure of the multimodal data, so as to optimize the information integration between the modes in the multimodal data and generate a manifold embedding intermediate result; A global information fusion mechanism is introduced for the manifold embedding intermediate result, and the manifold embedding intermediate result is further optimized by combining the global context information in the personalized problem understanding framework to generate a global optimized embedding graph; A weighted average strategy is introduced to assign different weights according to the importance of different data points in the global optimization embedding graph. A dimension selection technique is applied to select the dimension combination that is most representative of the multimodal data characteristics from the candidate dimensions generated by the global optimization embedding graph to generate a manifold embedding representation.

4. The method according to claim 2, characterized in that: The shared representation space is constructed by adopting multi-task learning technology, a plurality of interrelated subtasks are designed for the manifold embedding representation, and the interrelated subtasks are simultaneously trained in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a multi-task collaborative model, including: identifying natural clusters in the manifold embedding representation by a clustering method to obtain concentrated areas in the manifold embedding representation, evaluating the internal consistency and external separability of each cluster in the natural clusters, and generating a manifold embedding cluster analysis result; Adopting multi-task learning technology, constructing a shared representation space according to the manifold embedding cluster analysis result, designing multiple interrelated subtasks for the manifold embedding representation, and simultaneously training the interrelated subtasks in the shared representation space to improve the transfer efficiency of the personalized question understanding framework and generate a multi-task learning architecture; For the multi-task learning framework, a task relevance matrix is ​​introduced to quantify the task relevance in the multi-task learning framework, so as to enhance the synergy between subtasks in the multi-task learning framework and generate a task synergy optimization graph; In the shared representation space, an initial weight is assigned to each collaborative task in the task collaborative optimization graph, the initial weight is dynamically adjusted through a back-propagation algorithm, and an early stopping mechanism is introduced to prevent overfitting during the dynamic adjustment process, thereby generating a multi-task collaborative model.

5. The method according to claim 1, characterized in that: The recurrent neural network algorithm is used to recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, simulate the dynamic time change of the sequence information, so as to capture the dependency relationship between the sequence information, adopt a hierarchical clustering method, construct an organized information structure for the sequence information through a hierarchical aggregation method, and gradually merge the closest data points in the sequence information according to the similarity measurement in the organized information structure, provide deep knowledge mining, and generate an ordered information structure diagram, including: Applying time series feature extraction technology to extract time-dependent components in the cross-modal unified representation, analyzing the time pattern and periodicity features in the time-dependent components, and generating a time series feature set; Using a recurrent neural network algorithm, recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, input the internal state into the time series feature set, simulate the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, and generate a time series internal state diagram; Adopting a hierarchical clustering method, constructing an organized information structure for the sequence information and the internal state diagram of the time series by hierarchical aggregation, gradually merging the closest data points in the sequence information according to the similarity measurement in the organized information structure, providing deep knowledge mining, and generating a hierarchical clustering structure diagram; Identify the key nodes and paths in the hierarchical clustering structure diagram, apply a graph traversal algorithm, explore layer by layer from the root node to the leaf node in the key node, record the important transmission paths of ordered information in the layer-by-layer exploration process, and generate an ordered information structure diagram.

6. The method according to claim 5, characterized in that The recurrent neural network algorithm is used to recursively process the sequence information in the cross-modal unified representation, recursively update the internal state of the sequence information, input the internal state into the time series feature set, simulate the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, and generate a time series internal state diagram, including: Applying a sliding window mechanism, extracting local time series features from the time series feature set, combining time-frequency transformation technology, capturing global periodic patterns based on the local time series features, and generating a refined time series feature set; Using a recurrent neural network algorithm, recursively processing the sequence information in the cross-modal unified representation and the refined time series feature set, recursively updating the internal state of the sequence information, inputting the internal state into the time series feature set, simulating the dynamic time change of the sequence information, so as to capture the dependency relationship between the sequence information, and generating a recursively updated state record; Mapping the internal state of each time step in the recursively updated state record to the corresponding time point, and applying an anomaly detection algorithm to identify and mark potential abnormal mapping transitions to generate a time series state mapping table; The internal state information of all time steps in the time series state mapping table is integrated, and the mapping results in the time series state mapping table are converted into intuitive graphic representations by applying visualization technology to generate a time series internal state diagram.

7. The method according to claim 1, characterized in that The context-aware feedback mechanism is introduced to monitor the user's immediate reaction and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question answering solutions, including: Analyze the relationship between nodes and edges in the ordered information structure graph to identify the logical connection between different information elements in the ordered information structure graph, extract the information units and interaction modes represented by each node in the ordered information structure graph through pattern recognition technology, and generate a node and edge information table; Introducing a context-aware feedback mechanism to capture the user's immediate reaction and changes in the surrounding environment, and combining the node and edge information table to convert the immediate reaction and changes in the surrounding environment into computable feature vectors to generate a computable feature vector group; Dynamically adjust the answer strategy for the personalized question understanding framework, dynamically weight the computable feature vector group by performing context-dependent dynamic weight assignment on the ordered information structure graph, and generate a dynamic weighted feature vector; Evaluate the relevance and accuracy of the answer content in the answer strategy, optimize the evaluation result in combination with the dynamic weighted feature vector, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

8. An intelligent data analysis and knowledge question answering system based on a multimodal large model, characterized in that: include: A receiving module is used to receive and parse a multimodal query request from a user, perform semantic parsing and intent recognition technology on the multimodal query request, determine the core of the question and the required knowledge domain in the multimodal query request, and generate a personalized question understanding framework; the multimodal query request includes a combination of data in multiple forms such as text, image, audio and video; A processing module, used to use a manifold learning enhancement algorithm to perform manifold embedding processing on the multimodal data in the personalized question understanding framework, reveal the intrinsic geometric structure of the multimodal data, optimize the information integration between the modes in the multimodal data, use multi-task learning technology to construct a shared representation space, and simultaneously train multiple subtasks in the shared representation space to improve the transmission efficiency of the personalized question understanding framework and generate a cross-modal unified representation; The generating cross-modal unified representation process generates a manifold embedding representation, and the generating manifold embedding representation process generates a manifold embedding intermediate result; The generating manifold embedding intermediate result comprises: By principal component analysis, high-dimensional multimodal data in the consistent local neighborhood structure are mapped to a low-dimensional space, and a k-nearest neighbor graph is constructed to obtain similarity weights in the low-dimensional space to generate a local neighborhood structure consistency loss function; The local neighborhood structure consistency loss function is calculated using the following formula: ; in, is the local neighborhood structure consistency loss function; is the number of multimodal data samples; For sample of Neighbor set; For sample and The similarity weight between them is used to measure the correlation between samples in the local neighborhood; For sample and The square of the Euclidean distance between them is used to maintain the consistency of the local neighborhood structure; is the control coefficient of the logarithmic term; is the nonlinear adjustment coefficient, which is used to adjust the speed of exponential decay; Applying exponential decay functions and periodic functions to enhance the ability of the local neighborhood structure consistency loss function to capture complex geometric structures, adding regularization terms to prevent overfitting, and improving the generalization performance of the local neighborhood structure consistency loss function to generate a manifold embedding energy function; The manifold embedding energy function is calculated using the following formula: ; in, Embed energy function for manifold; is the exponential decay coefficient, which is used to adjust the strength of the nonlinear relationship; is the exponential decay rate coefficient; is the sine wave amplitude coefficient, used to introduce periodic changes; is the sine wave period adjustment coefficient; Is a balance parameter used to adjust the similarity weight The impact of For sample and The similarity weight between them; is a regularization term used to prevent overfitting; is the regularization coefficient, which is used to control the influence of the regularization term on the energy function; For low-dimensional representation and The square of the Euclidean distance between them is used to maintain the consistency of the internal geometric structure; and They are respectively and A low-dimensional representation of samples; Adopting the gradient descent method, iteratively updating the positions of the sample points in the manifold embedding energy function in the low-dimensional space to converge to the optimal solution, monitoring the energy change after each iteration during the iterative update process, and generating the manifold embedding intermediate result; A simulation module, for applying a recurrent neural network algorithm to recursively process the sequence information in the cross-modal unified representation, recursively updating the internal state of the sequence information, simulating the dynamic time change of the sequence information to capture the dependency relationship between the sequence information, using a hierarchical clustering method to construct an organized information structure for the sequence information through a hierarchical aggregation method, and gradually merging the closest data points in the sequence information according to the similarity measurement in the organized information structure, providing deep knowledge mining, and generating an ordered information structure diagram; The monitoring module is used to introduce a context-aware feedback mechanism, monitor the user's immediate reaction and environmental changes in real time, dynamically adjust the answer strategy for the personalized question understanding framework, evaluate and optimize the relevance and accuracy of the answer content in the answer strategy by performing context-related dynamic weight allocation on the ordered information structure diagram, record the user's immediate feedback on the answer content, and generate intelligent data analysis and knowledge question and answer solutions.

9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an intelligent data analysis and knowledge question-answering method based on a multimodal large model as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, an intelligent data analysis and knowledge question-answering method based on a multimodal large model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Time knowledge graph reasoning method and device based on multiple modes

    CN117668246A

  • Dynamic adaptation question answering system and method based on hierarchical structure and retrieval enhancement

    CN118193714A