Intelligent task matching method based on multi-modal data and RAG technology

By employing a smart task matching method that combines multimodal data conversion and RAG technology, the adaptability and accuracy issues of traditional single-modal processing methods are resolved. This approach enables efficient fusion of multimodal data and cross-task collaboration, thereby improving the adaptability and accuracy of task processing.

CN119903219BActive Publication Date: 2025-11-11XINGJI VALLEY (XIAN) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510269758.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-11-11
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

Traditional single-modal processing methods cannot meet the needs of multimodal data in complex scenarios, and have low adaptability and accuracy in task processing.

Method used

By collecting multimodal data, performing unified format conversion, using cross-modal collaborative attention mechanism for feature fusion, constructing RAG retrieval requirements, generating a database, and building a multi-learning framework to optimize feature sharing, intelligent task matching is performed.

Benefits of technology

It improves the adaptability and accuracy of intelligent task matching, realizes efficient fusion of multimodal data and cross-task collaboration, and enhances the effect of task processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903219B_ABST
    Figure CN119903219B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent task matching method based on multimodal data and RAG technology, belonging to the field of intelligent management technology. The method includes: collecting multimodal data, establishing modal data feature representations, and performing feature fusion based on a cross-modal collaborative attention mechanism; reading intelligent tasks, evaluating the fusion results, establishing RAG retrieval requirements, performing retrieval matching on an external database, and establishing a generated database; constructing and optimizing a multi-learning framework for intelligent tasks, and performing task processing for the corresponding intelligent tasks based on the converged multi-learning framework. This invention solves the technical problem that traditional single-modal processing methods in the prior art cannot meet the multimodal data requirements in complex scenarios, resulting in low adaptability and accuracy in task processing. It achieves the technical effect of improving the adaptability and accuracy of intelligent task matching through multimodal data fusion, RAG retrieval enhancement, and cross-task collaborative mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent management technology, specifically to an intelligent task matching method based on multimodal data and RAG technology. Background Technology

[0002] With the rapid development of multimedia technology, multimodal data (such as text, images, and audio) is increasingly widely used in various fields. How to efficiently process and fuse different types of data has become a significant challenge in the field of artificial intelligence. Traditional single-modal processing methods can no longer meet the needs of multimodal data in complex scenarios, especially when processing intelligent systems involving multiple tasks, resulting in low model adaptability and low task processing accuracy. Therefore, cross-modal feature fusion and task collaborative optimization have become crucial. Summary of the Invention

[0003] This application provides an intelligent task matching method based on multimodal data and RAG technology to address the technical problem that traditional single-modal processing methods in the prior art cannot meet the multimodal data requirements in complex scenarios, resulting in low adaptability and accuracy in task processing. RAG (Retrieval-Augmented Generation) technology is a technique that combines retrieval and generation, aiming to enhance the ability of large language models (LLMs) to handle knowledge-intensive tasks.

[0004] This application provides an intelligent task matching method based on multimodal data and RAG technology. The method includes: collecting multimodal data, establishing a multimodal dataset, and converting the multimodal dataset into a unified format using encoding technology to establish a multimodal data conversion result; establishing modal data feature representations based on the multimodal data conversion result, and performing feature fusion based on the modal data feature representations using a cross-modal collaborative attention mechanism to generate a fusion result of the multimodal data conversion result; reading intelligent tasks, performing data evaluation of the fusion result based on the intelligent tasks, and establishing data evaluation results, including data trust evaluation results and data richness evaluation results, wherein the intelligent tasks are at least two tasks; establishing RAG retrieval requirements based on the data evaluation results, performing retrieval matching on an external database, and establishing a generated database; using the fusion result and the generated database as basic data, synchronously executing the construction of a multi-learning framework for intelligent tasks, wherein the multi-learning framework is equipped with a cross-task collaboration module, and after feature sharing based on the cross-task collaboration module, optimizing the multi-learning framework based on a loss function; and performing task processing for the corresponding intelligent task based on the converged multi-learning framework.

[0005] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0006] This application provides an intelligent task matching method based on multimodal data and RAG technology, belonging to the field of intelligent management technology. It collects and transforms multimodal data, utilizes a cross-modal collaborative attention mechanism for feature fusion, constructs RAG retrieval requirements through intelligent task data evaluation, retrieves external databases to generate a new database, and builds a multi-task learning framework based on the fusion results and the generated database. This optimizes feature sharing and ultimately enables efficient intelligent task matching. It solves the technical problem that traditional single-modal processing methods in existing technologies cannot meet the multimodal data requirements of complex scenarios, resulting in low adaptability and accuracy in task processing. The method improves the adaptability and accuracy of intelligent task matching through multimodal data fusion, RAG retrieval enhancement, and cross-task collaborative mechanisms. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 A schematic diagram of the intelligent task matching method based on multimodal data and RAG technology provided in this application embodiment;

[0009] Figure 2 A schematic diagram illustrating the feature fusion based on modal data feature representation in the intelligent task matching method based on multimodal data and RAG technology provided in the embodiments of this application;

[0010] Figure 3 This is a schematic diagram illustrating the process of activating the cross-task collaboration module for feature sharing in the intelligent task matching method based on multimodal data and RAG technology provided in the embodiments of this application. Detailed Implementation

[0011] This application provides an intelligent task matching method based on multimodal data and RAG technology to solve the technical problem that traditional single-modal processing methods in the prior art cannot meet the multimodal data requirements in complex scenarios and have low adaptability and accuracy in task processing.

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0013] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0014] Examples, such as Figure 1 As shown, this application provides an intelligent task matching method based on multimodal data and RAG technology, the method comprising:

[0015] P10: Collect multimodal data, establish a multimodal dataset, and convert the multimodal dataset into a unified format using encoding techniques to establish the multimodal data conversion results.

[0016] Specifically, a comprehensive and standardized multimodal dataset is constructed to support subsequent intelligent task matching. First, multimodal data from diverse sources is collected extensively, including but not limited to text, images, audio, and potentially video and sensor data. This process relies on efficient data acquisition techniques and strategies to ensure data diversity and representativeness.

[0017] Subsequently, the collected multimodal data is organized into a structured dataset, namely a multimodal dataset. A multimodal dataset refers to a collection containing multiple data modalities, ensuring that the system can handle different types of information input. This step involves not only the physical storage of the data but also its logical organization and indexing to enable rapid access and processing later.

[0018] Because the data structures and characteristics of each modality differ (e.g., text is a linear structure, while images are two-dimensional pixel matrices), specific encoding techniques are needed to convert data from different modalities into a unified format. Encoding techniques can be based on neural networks (such as the Transformer encoder) or other deep learning models to represent text, images, audio, and other data as computationally usable embedding vectors. For text data, word embedding or sentence embedding techniques might be used to convert natural language text into vector representations in a high-dimensional space, capturing the semantic information within the text. For image data, feature maps extracted by convolutional neural networks (CNNs) might be used as the encoded representation of the image; these feature maps reflect the visual information within the image. Audio data might be converted into spectral features or time-series features using audio processing algorithms (such as MFCC, Mel Spectrogram, etc.).

[0019] Through the aforementioned encoding techniques, multimodal datasets were successfully converted into a unified format understandable by the model, i.e., the multimodal data conversion result. This process not only standardizes the data but also lays a solid foundation for subsequent operations such as feature fusion, data evaluation, and RAG retrieval. It is worth noting that the choice of encoding techniques should be optimized based on the specific task and data characteristics to ensure the accuracy and efficiency of the data conversion.

[0020] P20: Based on the multimodal data conversion results, establish modal data feature representations, and perform feature fusion based on modal data feature representations using a cross-modal collaborative attention mechanism to generate a fusion result of the multimodal data conversion results.

[0021] Furthermore, such as Figure 2 As shown, step P20 in this embodiment further includes:

[0022] P21: Place the modal data feature representation into the shared space, perform feature alignment loss analysis, and establish feature alignment results; P22: Perform granular segmentation of its own features, use the granular segmentation result as the smallest matching unit, perform association matching based on the feature alignment result, and establish a cross-modal collaborative attention mechanism after granular fusion of the association matching result; P23: Perform feature fusion of the modal data feature representation according to the cross-modal collaborative attention mechanism.

[0023] It should be understood that, firstly, a data feature representation for each modality is established based on the results of multimodal data transformation. This modal data feature representation is a high-dimensional representation obtained by extracting features from data of different modalities (such as text, images, and audio) using encoding techniques, ensuring that information from each modality can be effectively utilized in subsequent processing. For each modality, the system employs a specific feature extraction method to generate its data feature representation. For example, for text data, pre-trained natural language processing models (such as BERT, GPT, etc.) or word embedding techniques (such as Word2Vec, GloVe) are typically used to encode the text. Through these models, text can be represented as a high-dimensional vector, capturing the semantic and contextual information within the text. For instance, the sentence "I like apples" is encoded into an embedding vector using a BERT model; this vector represents the semantic information of the sentence. For image data, convolutional neural networks (CNNs) can be used to extract image features. CNNs extract low-level features (such as edges and textures) and high-level features (such as shape and objects) from images through multiple convolutional layers, pooling layers, and fully connected layers, ultimately representing the image as a feature vector. For example, an image containing apples can have its features extracted using a CNN model such as ResNet, transforming it into a vector representing the objects and their features contained in the image.

[0024] Next, these modal data feature representations are placed into a shared space, which can be achieved through a series of feature transformations and mapping operations. The aim is to eliminate the differences in representation between different modalities, enabling them to be compared and analyzed within the same framework. This shared space is a common representation space used for collaborative analysis of features from different modalities. Within this space, feature alignment loss analysis is performed. By calculating the similarity of features from different modalities, the alignment effect is analyzed, and feature alignment results are generated based on these results. This ensures that data from different modalities can be reasonably compared and fused within the same semantic or context for subsequent matching and fusion.

[0025] Next, granular segmentation of its own features is performed. Granular segmentation divides complex modal features into smaller feature units, ensuring that each smallest unit can be compared and matched more finely. These granular segmentation results are used as the smallest matching units for association matching based on feature alignment results. The goal of association matching is to find mutually related or complementary information by aligning features between different modalities. Subsequently, the association matching results are granularly fused to form a consistent representation, thereby establishing a cross-modal collaborative attention mechanism.

[0026] Cross-modal collaborative attention mechanism is one of the core technologies in this step. By focusing on important information in data from different modalities, it achieves feature fusion between modalities. The cross-modal collaborative attention mechanism is used to fuse the feature representations of modal data. This process involves deep interaction and fusion of features from different modalities within a shared space, generating a fusion result of multimodal data transformation. This ensures that data from different modalities can work collaboratively in a shared feature space, generating more accurate multimodal data fusion results. The fusion result not only includes the original information from each modality but also generates new and richer feature representations through cross-modal collaboration. These feature representations can more comprehensively reflect the essential characteristics of the data, providing strong support for subsequent intelligent task matching. In this way, the model can simultaneously utilize information from multiple modalities such as text, images, and audio, improving the accuracy and efficiency of intelligent task matching.

[0027] Furthermore, step P22 in the embodiments of this application also includes:

[0028] P22-1: Perform granular backtracking search based on the association matching results. The granular backtracking search is a search from the smallest granularity segmentation result to the largest granularity segmentation result. P22-2: Configure the association threshold and determine the termination granularity based on the granular backtracking search and the association threshold. P22-3: Establish an association interval based on the smallest granularity segmentation result and the termination granularity, and perform granular fusion based on the association matching results within the association interval.

[0029] Optionally, after completing the association matching, a granular backtracking search is performed. This process starts from the smallest granularity segmentation result and gradually searches towards larger granularity segmentation results. Granularity here refers to the coarseness of information units in the feature representation. The smallest granularity usually corresponds to the finest and most basic feature unit, while the largest granularity may contain broader and more comprehensive information. The purpose of granular backtracking search is to find the feature granularity combination that best suits the current task requirements, so as to more accurately capture and express the correlation information between data during feature fusion. Meanwhile, the implementation of granular backtracking search relies on efficient search algorithms, such as binary search, heuristic search, or graph-based search algorithms. These algorithms can minimize computational complexity and improve search efficiency while ensuring search accuracy.

[0030] During the granular backtracking search process, an association threshold needs to be configured. This threshold is used to evaluate the association strength between features at different granularities. When a certain granularity is reached, if its association strength is lower than the set threshold, it is considered that further increasing the granularity will no longer significantly improve the feature fusion effect, and the search terminates at this point; this granularity is the termination granularity. The setting of the association threshold is usually based on task requirements, data characteristics, and experimental experience. In practical applications, it may be necessary to find the most suitable threshold through multiple experiments and adjustments. Furthermore, some statistical methods (such as significance tests) can be used to assist in determining the threshold.

[0031] After determining the minimum granularity segmentation result and the termination granularity, an association interval is established between these two granularities. All granularity segmentation results within the association interval are considered valid matching units for subsequent granularity fusion operations. Further, granularity fusion is performed using the association matching results within the association interval. Granularity fusion refers to integrating and fusing different granularity segmentation results within the association interval to form the final cross-modal feature representation. Various methods can be employed, such as weighted averaging, max pooling, and attention mechanisms. These methods can flexibly adjust the contribution of different granularity features in the fusion result according to different task requirements and data characteristics. Furthermore, to ensure the stability and reliability of the fusion result, the fusion process needs to be thoroughly experimentally verified and optimized. This improves the accuracy and richness of the feature representation, providing a more solid foundation for subsequent intelligent task matching.

[0032] P30: Read the intelligent task, evaluate the data based on the fusion result of the intelligent task, and establish the data evaluation result, which includes the data trust evaluation result and the data richness evaluation result. The intelligent task is at least two tasks.

[0033] In one possible embodiment of this application, this step focuses on applying the fused multimodal data to a specific intelligent task and evaluating the effectiveness and quality of the data in task completion through data evaluation. First, the system reads and parses one or more intelligent tasks, which may involve multiple fields such as text classification, image recognition, and sentiment analysis, and each task has its specific goals and requirements.

[0034] Next, based on the specific needs of these intelligent tasks, the fusion results are evaluated. Data evaluation is a comprehensive process aimed at assessing the degree to which the data supports task completion from multiple dimensions, establishing data evaluation results, including data trustworthiness evaluation results and data richness evaluation results. Among these, the data trustworthiness evaluation results focus on the reliability and accuracy of the data, that is, the extent to which the data can truly reflect the actual situation or meet the task requirements. To assess data trustworthiness, various technical means can be used, such as data cleaning, outlier detection, and data consistency verification, to ensure the authenticity and accuracy of the data.

[0035] On the other hand, data richness evaluation focuses on assessing the completeness and diversity of the data. In intelligent task matching, abundant data often provides more comprehensive information support, thereby improving the accuracy and efficiency of task completion. To evaluate data richness, the source, type, quantity, and correlations and complementarities between the data can be analyzed. Through these analyses, the coverage and diversity of the dataset can be understood, thus determining whether it can meet the task requirements.

[0036] It is worth noting that since intelligent tasks may involve multiple fields and aspects, data evaluation results must also take into account the differences between different tasks. In practical applications, it may be necessary to customize different data evaluation indicators and methods for different tasks to ensure the accuracy and relevance of the evaluation. By establishing data evaluation results, important data support and assurance can be provided for subsequent intelligent task matching. This process not only helps improve the accuracy and efficiency of task completion but also provides feedback and guidance for continuous optimization of data quality.

[0037] P40: Based on the data evaluation results, establish RAG search requirements, perform search matching on external databases, and create a generated database.

[0038] Furthermore, step P40 in this embodiment of the application also includes:

[0039] P41: Call the data trust evaluation result and establish a trust verification retrieval requirement based on the data trust evaluation result; P42: Call the data richness evaluation result and establish a data supplementation retrieval requirement; P43: Establish a RAG retrieval requirement based on the trust verification retrieval requirement and the data supplementation retrieval requirement.

[0040] It should be understood that constructing and executing targeted RAG retrieval requirements based on data evaluation results aims to enhance and optimize the current dataset through retrieval matching from external databases, thereby generating a richer and more reliable generative database to enhance the processing capabilities of intelligent tasks.

[0041] First, the data trust evaluation results are retrieved, and trust verification retrieval requirements are constructed based on these results. The purpose of establishing trust verification retrieval requirements is to ensure that the data retrieved from external databases has a high degree of credibility with the existing fusion results, thereby guaranteeing the reliability and consistency of information. These requirements specifically target data segments marked as having low trust levels. By designing precise retrieval strategies, corresponding and more reliable data sources are retrieved from external databases to verify or replace existing data, thereby improving the overall trust level of the dataset. During the retrieval process, the retrieval strategy is adjusted based on the data trust level assessed in the evaluation, ensuring that the returned results meet the task's requirements for data trust level.

[0042] Next, the data richness evaluation results are invoked to establish data supplementation retrieval requirements. This requirement aims to further supplement the insufficient information in the fusion results through external databases, enhancing the richness and diversity of the data. For example, if the current task requires more background knowledge or contextual information, the system will retrieve relevant content from external databases through this retrieval process.

[0043] Finally, by combining the needs for trust verification and data supplementation in retrieval, a complete RAG retrieval requirement was constructed. This requirement allows the system to simultaneously meet the requirements for data credibility and data richness when performing external database matching. Through this dynamic retrieval and generation process, the system can obtain useful information from external databases, ultimately generating a richer, more reliable, and more diverse database, further improving the effectiveness of intelligent task matching.

[0044] In terms of technical support, to achieve efficient execution of RAG retrieval requirements, advanced retrieval technologies and algorithms are needed, such as Natural Language Processing (NLP) for understanding and parsing query requirements, distributed storage and computing technologies for processing large-scale data, and efficient indexing and ranking algorithms for improving retrieval speed and accuracy. Furthermore, a robust data management system is required to ensure accurate storage and rapid access to retrieval results.

[0045] P50: Using the fusion results and generated database as basic data, a multi-learning framework for synchronously executing intelligent tasks is built. The multi-learning framework is equipped with a cross-task collaboration module. After feature sharing is performed according to the cross-task collaboration module, the multi-learning framework is optimized based on the loss function.

[0046] Furthermore, such as Figure 3 As shown, step P50 in this embodiment further includes:

[0047] P51: Configure the task collaboration authentication channel. After performing sub-task collaboration authentication based on the task collaboration authentication channel, activate the cross-task collaboration module for feature sharing. The sub-task collaboration authentication includes: a: performing task relevance authentication of the sub-task and calculating the task relevance authentication result; b: calculating the feature sharing degree of the sub-task and generating the feature sharing degree calculation result; c: configuring the initial collaboration weight based on the task relevance authentication result and the feature sharing degree calculation result, and completing the sub-task collaboration authentication; P52: Activate the cross-task collaboration module for feature sharing based on the initial collaboration weight.

[0048] Optionally, the previously generated fusion results and the generated database can be used as the base data to build a multi-learning framework for intelligent tasks. The core of this multi-learning framework lies in its ability to process multiple tasks simultaneously, improving task processing efficiency by sharing underlying features. The framework also includes a cross-task collaboration module for feature sharing between different tasks. During task execution, the system continuously adjusts and improves the multi-learning framework by optimizing the loss function, ensuring optimal results for each task during collaborative processing.

[0049] Specifically, based on the fusion results and the generated database, a multi-learning framework for intelligent tasks is built. First, the system needs to clearly define the types and objectives of the intelligent tasks. For example, different intelligent tasks may include multimodal tasks such as text classification, image recognition, and speech processing. Next, the system further decomposes the tasks into multiple sub-tasks. Especially in the multi-task learning framework, tasks may have different focuses, such as classification, regression, and generation.

[0050] To support the construction of a multi-learning framework, the system needs to provide foundational data through fusion results and a generated database. During framework construction, a network architecture suitable for handling multiple tasks must be adopted. Common multi-task learning network architectures include a combination of a shared underlying network and task-specific branches. First, a shared underlying feature extraction network is established to extract common features from the fused multimodal data. This layer may be based on a deep learning model, such as a Convolutional Neural Network (CNN) for image processing, or a Transformer network for text processing. These networks are shared across multiple tasks, extracting common basic features and ensuring interoperability of information flow between tasks. Further, based on the shared feature layer, each intelligent task has its own task-specific network layer. These dedicated layers further process the shared features to generate task-specific output results. The dedicated layers for each task are designed according to the task's goals and requirements; for example, classification tasks might use fully connected layers, while generation tasks might use RNNs or Transformer decoders.

[0051] Furthermore, to ensure collaborative processing between multiple tasks, the system introduces a cross-task collaboration module. First, a task collaboration authentication channel is configured, through which sub-task collaboration authentication is performed. Sub-task collaboration authentication ensures sufficient relevance and reasonable sharing degree among tasks when sharing features. This sub-task collaboration authentication includes three key steps: First, task relevance authentication is performed on the sub-tasks, assessing their correlation by calculating similarity or dependency relationships, generating task relevance authentication results; second, feature sharing degree calculation is performed on the sub-tasks, evaluating the sharing potential and value of different tasks at the feature level, generating feature sharing degree calculation results; finally, based on the task relevance authentication results and feature sharing degree calculation results, initial collaboration weights are configured. For example, first, the system assigns a basic weight to each task group based on the task relevance results. Highly relevant tasks will receive higher weights because their synergistic effect is greater. Then, the system further adjusts the weights by referring to the feature sharing degree. If two tasks have high sharing degree, it means their features can be effectively utilized, and the weight will be further increased; if the sharing degree is low, the weight will be reduced accordingly. These weights reflect the importance and contribution of different tasks in the collaboration process, thus completing the sub-task collaboration certification.

[0052] After completing the collaborative authentication of the subtasks, the cross-task collaboration module is activated and begins feature sharing based on the initial collaboration weights. During this process, feature information from different tasks is circulated and integrated, enabling each task to extract useful information from other tasks and thus improve its performance. Feature sharing relies on advanced feature extraction and fusion techniques, which ensure that key information is retained while removing redundancy and noise, thereby improving feature utilization and effectiveness.

[0053] To further optimize the performance of the multi-learning framework, it is necessary to train and optimize the entire framework based on the loss function. The loss function is an important indicator that measures the difference between the model's predictions and the actual results. By continuously adjusting the model parameters to minimize the loss function value, the model can be made closer to the real situation, improving the accuracy and efficiency of intelligent tasks.

[0054] Furthermore, step P52 in this embodiment of the application also includes:

[0055] P52-1: Establish an adaptive weight update network for subtasks, as follows:

[0056] ; ;in, For subtasks Collaborative weights, For subtasks loss function, For model parameters, Characterization subtask The loss function relative to the model parameters gradient, These are parameters used to adjust and control the effect of the gradient on the weights. The normalization constant is This indicates the total number of subtasks. For subtask index; P52-2: Monitor feature sharing results by updating the network according to the adaptive weights, and update the collaborative weights with the monitoring results, and continue feature sharing with the updated collaborative weights.

[0057] It should be understood that further refining the working mechanism of the cross-task collaboration module, especially by introducing an adaptive weight update network to dynamically adjust the collaborative weights between subtasks, will improve the flexibility and adaptability of the multi-learning framework and ensure efficient feature sharing in different task scenarios.

[0058] Specifically, an adaptive weight update network for subtasks is first established. The core of this network lies in dynamically adjusting the collaborative weights of each subtask using the loss function of the subtask and its gradient information relative to the model parameters. The collaborative weights (denoted as ω) are calculated using the following formula: ; ;in, For subtasks Collaborative weights, For subtasks loss function, For model parameters, Characterization subtask The loss function relative to the model parameters gradient, These are parameters used to adjust and control the effect of the gradient on the weights. The normalization constant is This indicates the total number of subtasks. This serves as an index for subtasks. When β is large, subtasks with smaller gradients will receive higher weights, and vice versa. Finally, a normalization constant (i.e., the denominator of the sum of all subtask weights) ensures that the sum of all subtask weights is 1, maintaining the relativity and comparability of the weights.

[0059] Furthermore, the adaptive weight update network described above is used to monitor the feature sharing results. This monitoring process aims to evaluate the effectiveness of feature sharing in real time and dynamically adjust the collaborative weights based on the monitoring results. Specifically, if a subtask achieves a significant performance improvement through feature sharing, its collaborative weight may be increased accordingly to encourage more information sharing; conversely, if feature sharing does not help the subtask or even has a negative impact, its collaborative weight may be reduced to minimize unnecessary interference. In this way, it can be ensured that each subtask in the multi-learning framework can participate in the feature sharing process in an optimal manner, thereby maximizing overall performance. This not only improves the adaptability and robustness of the multi-learning framework but also provides more reliable and accurate data support for subsequent intelligent task matching.

[0060] P60: Perform task processing for the corresponding intelligent task based on the converged multi-learning framework.

[0061] Specifically, once the multi-learning framework reaches convergence through multiple iterations of training, the system can officially execute the corresponding intelligent task. Convergence means that the model's loss function gradually stabilizes during training, indicating that the model parameters have been adjusted to an optimal state, capable of effectively handling the input data in the task. In a multi-learning framework, this means that the collaborative work between the various sub-tasks has reached a relatively balanced state, and the effect of feature sharing is fully realized. A converged multi-learning framework possesses strong task processing capabilities, not only efficiently handling single tasks but also effectively operating in multi-task environments.

[0062] Subsequently, the converged multi-learning framework is used to execute the corresponding intelligent task. These intelligent tasks may include, but are not limited to, multiple fields such as image recognition, natural language processing, and speech recognition. During task processing, the multi-learning framework can fully utilize its internally constructed complex network structure and rich feature representation capabilities to perform deep analysis and processing of the input data, thereby extracting useful information and generating corresponding outputs. The key to this process lies in the fact that through previous multi-task learning and feature sharing, the multi-learning framework has accumulated rich experience and knowledge on multiple tasks. Therefore, during specific task execution, it can rely on this knowledge to complete task processing more quickly and accurately, thereby significantly improving the overall processing efficiency and accuracy of the intelligent system.

[0063] Furthermore, the embodiments of this application also include step P70, which further includes:

[0064] P71: Establish a feedback database, which is used to receive task processing results and extract processing feedback; P72: Establish a periodic optimization strategy based on the processing feedback, and use the periodic optimization strategy to perform adaptive optimization of multiple learning frameworks.

[0065] Optionally, to enhance the practicality and maintainability of the multi-learning framework, adaptive optimization can be achieved by introducing a feedback database and a periodic optimization strategy. First, a feedback database is established. The main function of the feedback database is to receive the results after each task and record the processing feedback related to that task. Processing feedback can include task execution efficiency, processing accuracy, error analysis, and other key performance indicators. Through this feedback, the system can gain a deeper understanding of the model's performance in actual task execution, thus providing a basis for subsequent optimization.

[0066] Next, based on the data in the feedback database, a periodic optimization strategy is formulated. This strategy is a regularly executed optimization mechanism designed to identify problems and shortcomings in the multi-learning framework's task processing by analyzing information in the feedback database, and to propose targeted improvement measures. Specifically, the strategy may include adjusting model parameter settings, optimizing feature extraction algorithms, and improving collaborative weight update mechanisms. By periodically executing these optimization measures, it is ensured that the multi-learning framework remains in a good working state, continuously improving its ability and efficiency in handling intelligent tasks.

[0067] In addition, to ensure the effective execution of the periodic optimization strategy, data mining techniques can be used to extract valuable information and patterns from the feedback database, machine learning algorithms can be used to automatically adjust and optimize model parameters based on this information, and automated testing tools can be used to ensure that no new errors or problems are introduced during the optimization process.

[0068] Furthermore, the embodiments of this application also include step P80, which further includes:

[0069] P81: Establish a domain database, configure incremental data through the domain database, and perform incremental learning of associated multi-learning frameworks; P82: Perform intelligent task matching for the corresponding domain based on the incremental learning results.

[0070] In one possible embodiment of this application, in-depth optimization and precise processing of a specific domain can be achieved by introducing a domain database and an incremental learning mechanism, enabling multiple learning frameworks to better adapt to the needs of different domains and improve the targeting and accuracy of task processing.

[0071] First, a domain database is established. This domain database is specifically designed to store and manage knowledge within a particular domain, containing various data samples, feature representations, label information, and domain-specific knowledge rules. By configuring incremental data into the domain database, new learning materials and challenges can be continuously provided to multiple learning frameworks, promoting their ongoing optimization and progress within the domain. Incremental data typically refers to data samples or information added to the existing dataset, reflecting the latest changes or trends within the domain.

[0072] Next, incremental learning within a multi-learning framework is performed using incremental data from the domain database. Incremental learning is a machine learning technique that allows a model to update and improve itself by learning from new data samples without retraining the entire dataset. In a multi-learning framework, incremental learning can be used for vertical optimization within a specific domain; that is, by learning from incremental data within that domain, the framework's processing power and performance in that domain are significantly improved. To achieve incremental learning, a series of advanced algorithms and techniques are required, such as online learning and transfer learning, to ensure that the model can efficiently absorb new knowledge and adapt to new environments.

[0073] Finally, intelligent task matching for the corresponding domain is performed based on the results of incremental learning. Since the multi-learning framework has been optimized for a specific domain through incremental learning, it is now able to more accurately understand and process tasks and data within that domain. During intelligent task matching, the multi-learning framework fully utilizes the knowledge and rules learned in the domain database to perform deep analysis and processing of the input data, generating output results that meet domain requirements. This process not only improves the accuracy and efficiency of task processing but also enhances the adaptability and flexibility of the multi-learning framework.

[0074] In summary, the embodiments of this application have at least the following technical effects:

[0075] This application collects and transforms multimodal data, utilizes a cross-modal collaborative attention mechanism for feature fusion, constructs RAG retrieval requirements through intelligent task data evaluation, retrieves external databases to generate a new database, and builds a multi-task learning framework based on the fusion results and the generated database to optimize feature sharing and ultimately efficiently perform intelligent task matching.

[0076] The technology has achieved the goal of improving the adaptability and accuracy of intelligent task matching through multimodal data fusion, RAG retrieval enhancement, and cross-task collaboration mechanisms.

[0077] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0078] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0079] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.

Claims

1. An intelligent task matching method based on multimodal data and RAG technology, characterized in that, The method includes: Collect multimodal data, establish a multimodal dataset, and convert the multimodal dataset into a unified format using encoding techniques to establish the multimodal data conversion result; Based on the multimodal data conversion results, a modal data feature representation is established, and feature fusion based on the modal data feature representation is performed based on a cross-modal collaborative attention mechanism to generate a fusion result of the multimodal data conversion results; Read the intelligent task, evaluate the data based on the fusion result of the intelligent task, and establish the data evaluation result, which includes the data trust evaluation result and the data richness evaluation result. The intelligent task is at least two tasks. Based on the data evaluation results, establish RAG search requirements, perform search matching on external databases, and create a generated database; Using the fusion result and the generated database as the basic data, a multi-learning framework for synchronously executing intelligent tasks is built. The multi-learning framework is equipped with a cross-task collaboration module. After feature sharing is performed according to the cross-task collaboration module, the multi-learning framework is optimized based on the loss function. Task processing is performed based on the converged multi-learning framework to execute the corresponding intelligent task. The feature fusion based on modal data feature representation using the cross-modal collaborative attention mechanism to generate multimodal data transformation results also includes: The modal data feature representations are placed into a shared space, feature alignment loss analysis is performed, and feature alignment results are established. Perform granular segmentation of its own features, use the granular segmentation result as the smallest matching unit, perform association matching based on feature alignment result, and establish a cross-modal collaborative attention mechanism after granular fusion of association matching result; Feature fusion of modal data feature representation based on the cross-modal collaborative attention mechanism.

2. The intelligent task matching method based on multimodal data and RAG technology as described in claim 1, characterized in that, The step of optimizing the multi-learning framework based on a loss function after feature sharing according to the cross-task collaboration module further includes: Configure a task collaboration authentication channel. After performing sub-task collaboration authentication based on the task collaboration authentication channel, activate the cross-task collaboration module for feature sharing. The sub-task collaboration authentication includes: a: Perform task relevance authentication for subtasks and calculate the task relevance authentication result; b: Calculate the feature sharing degree for the subtask and generate the feature sharing degree calculation result; c: Configure initial collaborative weights based on the task relevance authentication results and feature sharing degree calculation results to complete sub-task collaborative authentication; The cross-task collaboration module is activated to share features based on the initial collaboration weights.

3. The intelligent task matching method based on multimodal data and RAG technology as described in claim 2, characterized in that, The activation of the cross-task collaboration module for feature sharing based on the initial collaboration weights further includes: Establish an adaptive weight update network for subtasks as follows: ; ; in, For subtasks Collaborative weights, For subtasks loss function, For model parameters, Characterization subtask The loss function relative to the model parameters gradient, These are parameters used to adjust and control the effect of the gradient on the weights. The normalization constant is This indicates the total number of subtasks. Index for subtasks; The adaptive weight update network is used to monitor the feature sharing results, and the collaborative weights are updated based on the monitoring results. Feature sharing continues with the updated collaborative weights.

4. The intelligent task matching method based on multimodal data and RAG technology as described in claim 1, characterized in that, The granular fusion of the association matching results also includes: Granularity backtracking search is performed based on the association matching results. The granularity backtracking search is a search from the smallest granularity segmentation result to the largest granularity segmentation result. Configure an association threshold, and determine the termination granularity based on the granularity backtracking search and the association threshold; Based on the minimum granularity segmentation result and the termination granularity, an association interval is established, and granularity fusion is performed using the association matching results within the association interval.

5. The intelligent task matching method based on multimodal data and RAG technology as described in claim 1, characterized in that, The step of establishing RAG search requirements based on the data evaluation results also includes: The data trust evaluation result is invoked, and a trust verification retrieval requirement is established based on the data trust evaluation result; Use the data richness evaluation results to establish data supplementation retrieval requirements; RAG search requirements are established based on the trust verification search requirements and data supplementation search requirements.

6. The intelligent task matching method based on multimodal data and RAG technology as described in claim 1, characterized in that, The method further includes: Establish a feedback database, which is used to receive task processing results and extract processing feedback; A periodic optimization strategy is established based on the processing feedback, and adaptive optimization of multiple learning frameworks is performed using the periodic optimization strategy.

7. The intelligent task matching method based on multimodal data and RAG technology as described in claim 1, characterized in that, The method further includes: Establish a domain database, configure incremental data through the domain database, and perform incremental learning of associated multi-learning frameworks; Based on the incremental learning results, perform intelligent task matching in the corresponding domain.

Citation Information

Patent Citations

  • Multi-modal fusion emotion recognition system and method based on multi-task learning and attention mechanism and experimental evaluation method

    CN113420807A

  • Electric power cross-modal knowledge fusion multi-agent cooperative processing method and system

    CN119477235A