A method and system for constructing a machine language large model

Through intelligent semantic association network, cross-modal attention mechanism and adversarial generation network, combined with multi-level feedback system and task adaptive module, the problems of insufficient deep semantic association and poor adaptability in multi-modal data processing in the prior art are solved, and more efficient multi-modal data understanding and dynamic adaptation are achieved.

CN119721118BActive Publication Date: 2025-05-30HUAQING WEIYANG (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510220620.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The prior art fails to fully mine deep semantic associations between text and visual or auditory data when processing multimodal data, and the model is less adaptable in the face of new tasks or environmental changes.

Method used

An intelligent semantic association network is adopted to combine cross-modal attention mechanisms and adversarial generation networks to enhance the association between text and visual or auditory data, and dynamically adjust model parameters through a multi-level feedback system and task adaptive module to generate an adaptive learning rate adjustment strategy.

Benefits of technology

The model's understanding and dynamic adaptability of multimodal data is significantly improved, ensuring that the model maintains high performance and flexibility in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721118B_ABST
    Figure CN119721118B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for constructing a machine language large model, which receives user input data, defines an intelligent semantic association network, and extracts semantic associations between text, visual, or auditory data; uses a cross-modal attention mechanism and a generative adversarial network to enhance the multi-modal data relationships in these associations, forming a multi-modal perception mechanism; based on this mechanism, constructs a multi-level feedback system to process feedback signals from different levels and generate an optimized prediction path; combines a task adaptive module and the optimized prediction path to dynamically adjust key parameters and generate an adaptive learning rate adjustment strategy; constructs a machine language large model based on an incremental context expansion module, an intelligent semantic association network, a multi-modal perception mechanism, a multi-level feedback system, and a task adaptive module. The present application improves the model's understanding and processing capabilities for complex multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the fields of machine learning and natural language technologies, and in particular, to a method and system for constructing a large machine language model. Background Art

[0002] In today's field of artificial intelligence, the ability to process multimodal data is crucial for achieving the understanding and interaction of complex tasks. Current technical solutions are widely applied in multiple fields such as intelligent assistants, autonomous driving, and medical diagnosis. These systems usually adopt independent or simple fusion methods to process data of different modalities. Specifically, existing methods first process data of each modality separately, and then fuse the results in some way at a later stage, or perform a preliminary fusion of data of different modalities during the feature extraction stage. Although these methods can achieve certain effects in specific application scenarios, they fail to fully explore the deep semantic associations between modalities, restricting the comprehensive utilization of information.

[0003] Existing multimodal processing solutions usually adopt independent or simple fusion methods to process data of different modalities. Some methods process data of each modality separately and then fuse the results in some way at a later stage. Although this method is simple and direct, it fails to fully utilize the associations between modalities. Another common method is to perform a preliminary fusion of data of different modalities during the feature extraction stage, but this shallow fusion lacks the exploration of deep semantic associations and cannot effectively capture the complex relationships between modalities. In addition, some models use fixed architectures and parameter settings, and although they can perform well on specific tasks, they are less adaptable when facing new tasks or environmental changes. Although these existing solutions have made progress in some aspects, they have not comprehensively solved the challenges in multimodal data processing, especially in terms of dynamic adaptability and deep semantic associations, there are still deficiencies.

[0004] Existing solutions have obvious limitations in processing multimodal data, especially in terms of deep semantic associations and dynamic adaptability. First of all, most existing methods fail to fully explore the deep semantic associations between text and related visual or auditory data, resulting in low information utilization rate and affecting the overall performance of the model. Secondly, many models rely on fixed parameter configurations and are difficult to dynamically adjust according to task changes, which restricts the generalization ability and adaptability of the model. In addition, most existing systems lack an effective multi-level feedback mechanism and cannot optimize the prediction path in real time, thus restricting the performance improvement of the model in complex environments. Summary of the Invention

[0005] Embodiments of the present application provide a method and system for constructing a large machine language model to solve the problem of low understanding and processing ability of complex multimodal data in the prior art.

[0006] In a first aspect, an embodiment of the present application provides a method for constructing a machine language large model, including:

[0007] Receiving user input data, defining an intelligent semantic association network, and obtaining semantic associations based on the intelligent semantic association network in combination with the user input data; using a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic associations and visual or auditory data related to the text, and determining a multi-modal perception mechanism;

[0008] Based on the multi-modal perception mechanism, constructing a multi-level feedback system, and using the multi-level feedback system to process feedback signals from different levels to generate an optimized prediction path;

[0009] Based on the multi-level feedback system, constructing a task adaptive module using a meta-learning framework and a reinforcement learning algorithm, and based on the task adaptive module and the optimized prediction path, adjusting key parameters in the task adaptive module to generate an adaptive learning rate adjustment strategy;

[0010] Based on the optimized prediction path and the adaptive learning rate adjustment strategy, constructing an incremental context expansion module using a structured representation learning technique based on a graph neural network, and constructing a machine language large model based on the incremental context expansion module, the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptive module.

[0011] Optionally, the constructing a task adaptive module using a meta-learning framework and a reinforcement learning algorithm based on the multi-level feedback system, and based on the task adaptive module and the optimized prediction path, adjusting key parameters in the task adaptive module to generate an adaptive learning rate adjustment strategy includes:

[0012] Using a meta-learning framework and a reinforcement learning algorithm to construct an initial task adaptive module based on the multi-level feedback system, and optimizing the conversion efficiency of the initial task adaptive module between different tasks through a Q-learning algorithm and a policy gradient method;

[0013] Based on the optimized prediction path, using a variational autoencoder to analyze the change trend from the data distribution of feedback signals at different levels in the multi-level feedback system, and dynamically adjusting key parameters in the initial task adaptive module through Bayesian optimization to generate a target task adaptive module;

[0014] Using a context awareness mechanism to enable the target task adaptive module to adjust the learning strategy according to the background information of different tasks to generate an adaptive learning rate adjustment strategy.

[0015] Optionally, using the context awareness mechanism, the target task adaptive module adjusts the learning strategy according to the background information of different tasks to generate an adaptive learning rate adjustment strategy, including:

[0016] Using the context awareness mechanism, combining the graph neural network and the attention mechanism, analyze the key features from the background information of different tasks in the multi-level feedback system, where the key features include task type, data distribution characteristics, and historical performance;

[0017] Based on the key features, combining Bayesian optimization and Gaussian process regression, dynamically adjust the learning rate in the target task adaptive module to generate a preliminary adaptive learning rate adjustment strategy;

[0018] Using a time series prediction model, optimize the preliminary adaptive learning rate adjustment strategy through reinforcement learning to generate an optimized learning rate adjustment strategy;

[0019] Combining genetic algorithms and differential evolution algorithms, search for the optimal learning rate adjustment strategy from the optimized learning rate adjustment strategy;

[0020] Introduce an active learning mechanism, select target data from the optimal learning rate adjustment strategy to generate an adaptive learning rate adjustment strategy.

[0021] Optionally, using the context awareness mechanism, combining the graph neural network and the attention mechanism, analyze the key features from the background information of different tasks in the multi-level feedback system, including:

[0022] Using data mining techniques, collect and preprocess the background information of different tasks in the multi-level feedback system to obtain the task background data of the multi-level feedback system;

[0023] Introduce a context awareness mechanism, combine principal component analysis and factor analysis, perform preliminary screening and dimensionality reduction processing on the task background data to obtain the target task background information;

[0024] Apply graph neural network and complex network analysis to analyze the complex relationships between tasks from the target task background information to construct a task relationship graph;

[0025] Apply the attention mechanism, combine with deep reinforcement learning, highlight the important features in the task relationship graph to obtain the key feature weight assignment result;

[0026] Analyze the task relationship graph and the key feature weight assignment result to extract the key features.

[0027] Optionally, based on the optimized prediction path, using a variational autoencoder, analyze the changing trends from the data distributions of feedback signals at different levels in the multi-level feedback system, and dynamically adjust the key parameters in the initial task adaptation module through Bayesian optimization to generate a target task adaptation module, including:

[0028] According to the optimized prediction path, use a variational autoencoder and statistical analysis tools to model the data distributions of feedback signals at different levels in the multi-level feedback system, and analyze potential patterns and changing rules from the feedback information to generate a changing trend of the data distribution;

[0029] Based on the changing trend of the data distribution, dynamically adjust the key parameters in the initial task adaptation module through Bayesian optimization, Gaussian process regression, and gradient boosting trees to generate a target task adaptation module.

[0030] Optionally, use a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic association and the visual or auditory data related to the text, and determine a multi-modal perception mechanism, including:

[0031] Use a hierarchical graph neural network to define an initial intelligent semantic association network, and analyze an initial semantic association from the internal connection between the text content of the initial intelligent semantic association network and the visual or auditory data associated with the text content;

[0032] Based on the initial semantic association, introduce a cross-modal attention mechanism, and combine a bidirectional long short-term memory network to perform joint encoding processing on the text and the corresponding visual or auditory data in the initial semantic association to obtain a multi-modal information representation;

[0033] According to the multi-modal information representation, construct a generative adversarial network, and analyze key data from the adversarial training process between the generator and the discriminator in the generative adversarial network;

[0034] Based on the key data, combine a variational autoencoder to optimize the initial intelligent semantic association network to obtain a target intelligent semantic association network, and determine a multi-modal perception mechanism based on the target intelligent semantic association network.

[0035] Optionally, based on the optimized prediction path and the adaptive learning rate adjustment strategy, use a structured representation learning technique based on a graph neural network to construct an incremental context expansion module, including:

[0036] Use the optimized prediction path and the adaptive learning rate adjustment strategy, and combine time series analysis and Bayesian optimization to determine an incremental update rule;

[0037] Based on the incremental update rule, using the structured representation learning technology based on graph neural network, combined with the hierarchical clustering algorithm, model the multimodal data in the multimodal perception mechanism to obtain the structured representation of multimodal data;

[0038] According to the structured representation of multimodal data, combined with the dynamic graph convolutional network and variational autoencoder, encode the data obtained from the multi-level feedback system, and input the data into the structured representation of multimodal data to output the updated context representation;

[0039] Apply the attention mechanism, combined with deep reinforcement learning, to highlight the key features in the updated context representation to obtain the optimized context representation;

[0040] Combine reinforcement learning and genetic algorithm to optimize and adjust the parameters in the incremental update process of the task adaptation module to generate an optimized parameter configuration;

[0041] Based on the optimized context representation and the optimized parameter configuration, generate an incremental context expansion module.

[0042] In a second aspect, an embodiment of the present application provides a machine language large model construction system, including:

[0043] A receiving module, configured to receive user input data, define an intelligent semantic association network, and based on the intelligent semantic association network and the user input data, obtain a semantic association;

[0044] An enhancement module, configured to use a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic association and the visual or auditory data related to the text, and determine a multimodal perception mechanism;

[0045] A construction module, configured to construct a multi-level feedback system based on the multimodal perception mechanism, and use the multi-level feedback system to process feedback signals from different levels to generate an optimized prediction path;

[0046] An adjustment module, configured to construct a task adaptation module based on the multi-level feedback system, using a meta-learning framework and a reinforcement learning algorithm, and based on the task adaptation module and the optimized prediction path, adjust the key parameters in the task adaptation module to generate an adaptive learning rate adjustment strategy;

[0047] A utilization module is used to construct an incremental context expansion module by using the structured representation learning technology based on graph neural network based on the optimized predicted path and the adaptive learning rate adjustment strategy. A machine language large model is constructed based on the incremental context expansion module, the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptation module. In a third aspect, an embodiment of the present application provides a computing device, including a processor and a memory. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for constructing a machine language large model according to any one of the first aspects.

[0048] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method for constructing a machine language large model according to any one of the first aspects is implemented.

[0049] In the embodiment of the present application, user input data is received, and an intelligent semantic association network is defined. Based on the intelligent semantic association network and the user input data, semantic associations are obtained; the cross-modal attention mechanism and the adversarial generation network are used to enhance the relationship between the text in the semantic associations and the visual or auditory data related to the text, and a multi-modal perception mechanism is determined; based on the multi-modal perception mechanism, a multi-level feedback system is constructed, and the multi-level feedback system is used to process feedback signals from different levels to generate an optimized predicted path; based on the multi-level feedback system, a task adaptation module is constructed by using a meta-learning framework and a reinforcement learning algorithm. Based on the task adaptation module and the optimized predicted path, key parameters in the task adaptation module are adjusted to generate an adaptive learning rate adjustment strategy; based on the optimized predicted path and the adaptive learning rate adjustment strategy, an incremental context expansion module is constructed by using the structured representation learning technology based on graph neural network. A machine language large model is constructed based on the incremental context expansion module, the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptation module.

[0050] The technical solution of the present application has the following beneficial effects:

[0051] The intelligent semantic association network of this application effectively captures the deep semantic connections between text and relevant visual or auditory data, improving the model's ability to understand multimodal data. Through the cross-modal attention mechanism and the adversarial generation network, the association between text and other modal data is enhanced, enabling the model to more comprehensively understand complex multimodal inputs. The multi-level feedback system processes different levels of feedback signals in real time, generating optimized prediction paths to ensure that the model can maintain high performance in different tasks. The task adaptation module combines the optimized prediction path, dynamically adjusts key parameters, and generates an adaptive learning rate adjustment strategy, enabling the model to quickly adapt and continuously optimize in a changing environment. Based on the incremental context expansion module, the model can gradually introduce new information, continuously improve while maintaining existing performance, ensuring long-term effectiveness and flexibility. The overall architecture integrates the intelligent semantic association network, the multimodal perception mechanism, the multi-level feedback system, and the task adaptation module, constructing a powerful large machine language model with the ability to efficiently process complex tasks.

[0052] Furthermore, the embodiment of this application also constructs a task adaptation module through the multi-level feedback system in combination with the meta-learning framework and the reinforcement learning algorithm, and adjusts key parameters based on this module and the optimized prediction path to generate an adaptive learning rate adjustment strategy. Specifically, first, using the meta-learning framework and the reinforcement learning algorithm, an initial task adaptation module is constructed based on the multi-level feedback system, and the conversion efficiency of this module between different tasks is optimized through the Q-learning algorithm and the policy gradient method; then, based on the optimized prediction path, the variational autoencoder is used to analyze the data distribution change trend of different levels of feedback signals in the multi-level feedback system, and the key parameters in the initial task adaptation module are dynamically adjusted through Bayesian optimization to generate the final target task adaptation module; finally, the context-aware mechanism is used to enable the target task adaptation module to adjust the learning strategy according to different task background information, thereby generating an adaptive learning rate adjustment strategy.

[0053] Through the above method, not only the task adaptation ability and learning efficiency of the model are significantly improved, but also it is ensured that the model can continuously optimize when facing new tasks and environmental changes. The task adaptation module constructed by the meta-learning framework and reinforcement learning algorithm can efficiently switch between different tasks, improving the flexibility and generalization ability of the model. The application of the variational autoencoder combined with Bayesian optimization enables the model to extract valuable changing trends from the complex data distribution of the multi-level feedback system, dynamically adjust key parameters, and further enhance the adaptability of the model to different task backgrounds. In addition, the application of the context-aware mechanism enables the model to flexibly adjust the learning strategy according to specific task background information, generate the optimal learning rate adjustment strategy, thus ensuring the high-performance performance and fast convergence of the model in various application scenarios. In summary, the present method effectively solves the problems existing in the existing solutions, such as insufficient deep semantic association, static parameter configuration, and imperfect feedback mechanism, and provides a more intelligent and efficient multi-modal processing solution.

[0054] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. Brief Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a flowchart of a method for constructing a machine language large model provided by an embodiment of the present application;

[0057] Figure 2 It is a schematic structural diagram of a system for constructing a machine language large model provided by an embodiment of the present application;

[0058] Figure 3 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Detailed Embodiments

[0059] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0060] In some processes described in the specification, claims, and the above-mentioned drawings of this application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear herein or in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations can be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0062] Figure 1 The flow chart of a method for constructing a machine language large model is provided for an embodiment of the present application, as Figure 1 shown. The method includes:

[0063] Step 101: Receive user input data, define an intelligent semantic association network, and based on the intelligent semantic association network and the user input data, obtain semantic associations;

[0064] In this step, multi-modal input data from the user is received, and these data can be in the form of text, images, or audio, etc. The intelligent semantic association network is a complex neural network architecture designed to capture the deep semantic relationships between various modal data such as text, images, and audio. Through multi-level feature extraction and association analysis, the network can identify the implicit connections between different modal data, thereby providing rich semantic information for subsequent processing. For example, in natural language processing, it can identify the keywords in a sentence and match them with the key elements in relevant image or audio segments to enhance the understanding of multi-modal data.

[0065] In actual operation, after receiving the user input, the system preprocesses the multi-modal data through a preprocessing module, such as text tokenization, image feature extraction, and audio transcription. Then, an intelligent semantic association network is defined, and the implicit connections between different modal data are identified through multi-level feature extraction and association analysis to generate specific semantic associations. These associations not only include direct mapping relationships but also cover indirect semantic connections, providing a solid foundation for subsequent steps.

[0066] For example, in a smart assistant application, users may interact with the system through voice commands, text input, or uploading pictures. For example, a user uploads a photo with the label "beach" and asks about nearby restaurants. The system first receives this photo and label information. After preprocessing, it uses an intelligent semantic association network to identify visual elements such as beaches and waves in the photo and establish an association with the "beach" label, generating an initial semantic association to provide a basis for the next step of multimodal perception.

[0067] Step 102: Use a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic association and the visual or auditory data related to the text, and determine a multimodal perception mechanism;

[0068] In this step, the cross-modal attention mechanism dynamically adjusts the importance of different modal data in the fusion process by introducing attention weights, enabling the model to focus on the most relevant parts. The generative adversarial network is used to generate high-quality synthetic data to enhance the model's understanding of the relationship between multimodal data. This method can not only improve the model's processing ability for existing data but also enrich the training samples by generating new data, enhancing the model's generalization ability.

[0069] In actual operation, using the cross-modal attention mechanism, the model calculates the attention scores between the text and other modal data, and these scores reflect the correlation of different modal data. Then, the generative adversarial network is used to generate new multimodal data pairs, and after training, these data pairs further enhance the model's understanding of multimodal relationships. Finally, through this method, a multimodal perception mechanism is determined, enabling the model to be more sensitive and accurate when processing multimodal data.

[0070] For example, continuing with the above example of the smart assistant, when the user asks about nearby restaurants, the system uses the cross-modal attention mechanism to identify the key elements in the photo (such as beaches and waves) and establish a stronger association with the text label "beach". At the same time, the generative adversarial network generates more similar multimodal data pairs, making the model perform more robustly and accurately when facing similar queries. This not only improves the system's response speed but also enhances user satisfaction.

[0071] Step 103: Based on the multimodal perception mechanism, construct a multi-level feedback system, and use the multi-level feedback system to process feedback signals from different levels to generate an optimized prediction path;

[0072] In this step, the multi-level feedback system is a hierarchical structure, with each layer responsible for processing specific types of feedback signals, from low-level raw data to high-level task results. Through this hierarchical processing, the system can gradually refine and optimize its prediction path, ensuring that each decision is based on comprehensive and accurate information. The multi-level feedback system can not only adjust the behavior of the model in real time but also record historical feedback to provide a basis for future improvement.

[0073] In actual operation, based on the multi-modal perception mechanism, a multi-level feedback system is constructed. This system is divided into multiple levels, and each level processes feedback signals at different levels. For example, the lowest level processes the raw multi-modal data, the middle level processes the preliminary feature extraction results, and the highest level processes the final task output. Through layer-by-layer processing and feedback, the system can continuously optimize its prediction path to ensure that each prediction is optimal. In addition, the system will also adjust parameters according to historical feedback to further improve the accuracy of prediction.

[0074] For example, in the intelligent assistant application, the multi-level feedback system first receives the photos and voice commands uploaded by the user as the raw data at the lowest level. Next, the system processes this data at the middle level and extracts key features, such as the names of dishes on the menu and the word "restaurant" in the voice. Finally, at the highest level, the system synthesizes this information to generate an optimized prediction path, that is, a list of recommended nearby restaurants. If the user is not satisfied with the recommendation, the system will collect feedback and make corresponding adjustments during the next query, gradually optimizing the recommendation result to provide a better user experience.

[0075] Step 104: Based on the multi-level feedback system, use the meta-learning framework and reinforcement learning algorithm to construct a task adaptation module. Based on the task adaptation module and the optimized prediction path, adjust the key parameters in the task adaptation module to generate an adaptive learning rate adjustment strategy;

[0076] In this step, the task adaptation module enables the model to quickly switch between different tasks and efficiently adapt to new tasks through the meta-learning framework and reinforcement learning algorithm. The meta-learning framework allows the model to learn experience from previous tasks, while the reinforcement learning algorithm optimizes the behavior of the model through a reward mechanism. By dynamically adjusting the key parameters, the task adaptation module can generate an adaptive learning rate adjustment strategy to ensure that the model can maintain the best performance in different tasks and environments.

[0077] In actual operation, based on the optimized predicted path provided by the multi-level feedback system, a task adaptive module is constructed using a meta-learning framework and a reinforcement learning algorithm. This module optimizes the conversion efficiency between different tasks through the Q-learning algorithm and the policy gradient method, ensuring that the model can quickly adapt to new tasks. At the same time, the key parameters in the initial task adaptive module are dynamically adjusted through Bayesian optimization to generate the final target task adaptive module. This process enables the model to flexibly adjust the learning strategy according to the current task context and generate the optimal learning rate adjustment strategy.

[0078] For example, in a smart assistant application, the task adaptive module can quickly adjust its parameter configuration according to the user's historical interaction records and the current query content to adapt to different query types. For example, when the user frequently queries restaurant information, the module will optimize the parameters related to restaurants to improve the query efficiency. When the user starts asking about tourist attractions, the module can quickly switch to the parameter configuration suitable for the tourist scenario. In this way, the system can not only give accurate recommendations in a short time but also continuously optimize itself as the user's behavior changes to provide more personalized services.

[0079] Step 105: Based on the optimized predicted path and the adaptive learning rate adjustment strategy, use the structured representation learning technology based on graph neural networks to construct an incremental context expansion module, and construct a large machine language model based on the incremental context expansion module, the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptive module.

[0080] In this step, the incremental context expansion module realizes structured representation learning through graph neural networks, enabling the model to gradually introduce new information without affecting the existing performance. This module can dynamically expand the context range according to the optimized predicted path and the adaptive learning rate adjustment strategy, ensuring that the model remains efficient when processing long sequences or multi-step reasoning tasks. Finally, a complete large machine language model is constructed by combining the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptive module.

[0081] In actual operation, based on the optimized predicted path and the adaptive learning rate adjustment strategy, use the structured representation learning technology of graph neural networks to construct an incremental context expansion module. This module expands the context range by gradually introducing new information, ensuring that the model remains efficient when processing complex tasks. By combining the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptive module, a powerful large machine language model is finally constructed. This model can perform excellently in various application scenarios and has efficient adaptive capabilities and incremental learning characteristics.

[0082] For example, in a smart assistant application, the incremental context expansion module can gradually expand the context scope when the user continuously asks multiple questions, ensuring that each answer is based on the latest information. For example, the user first asks about nearby restaurants, then asks about the specific dishes of a certain restaurant, and finally asks about the business hours of this restaurant. The module will gradually introduce this new information and expand the context scope to ensure that each answer is accurate. At the same time, other components such as the intelligent semantic association network and multi-modal perception mechanism work together to jointly improve the overall performance of the system and provide users with a seamless interaction experience.

[0083] Through the above five steps, this application constructs a highly intelligent machine language large model, significantly improving the model's understanding and processing capabilities for multi-modal data. The intelligent semantic association network and cross-modal attention mechanism ensure the mining of deep semantic associations. The multi-level feedback system and task adaptation module achieve efficient dynamic adjustment and optimization, while the incremental context expansion module guarantees the flexibility and high performance of the model when dealing with complex tasks. Finally, the model not only performs well in the initial performance but also can continuously improve when facing new tasks and environmental changes, and is applicable to a wide range of natural language processing and other multi-modal application fields, greatly enhancing the user experience and system efficiency.

[0084] To solve the problems that the task adaptation module in the existing methods is not flexible enough in adjustment and the learning rate adjustment strategy lacks dynamics, in one or more of the above embodiments, in step 104, based on the multi-level feedback system, a task adaptation module is constructed using a meta-learning framework and a reinforcement learning algorithm, and based on the task adaptation module and the optimized prediction path, the key parameters in the task adaptation module are adjusted to generate an adaptive learning rate adjustment strategy, including:

[0085] Using a meta-learning framework and a reinforcement learning algorithm, an initial task adaptation module is constructed based on the multi-level feedback system, and the conversion efficiency of the initial task adaptation module between different tasks is optimized through the Q-learning algorithm and the policy gradient method; based on the optimized prediction path, using a variational autoencoder, the change trend is analyzed from the data distribution of different-level feedback signals in the multi-level feedback system, and the key parameters in the initial task adaptation module are dynamically adjusted through Bayesian optimization to generate a target task adaptation module; using a context awareness mechanism, the target task adaptation module adjusts the learning strategy according to the background information of different tasks to generate an adaptive learning rate adjustment strategy.

[0086] In this embodiment, the construction of the task adaptation module not only relies on traditional static parameter configuration, but also introduces a meta-learning framework and a reinforcement learning algorithm. The meta-learning framework allows the model to quickly learn experience from previous tasks and apply it to new tasks, thereby improving the efficiency of task switching. The reinforcement learning algorithm optimizes the behavior of the model through a reward mechanism to ensure its efficient switching between different tasks. In addition, the variational autoencoder is used to analyze the data distribution change trend of the multi-level feedback system, and Bayesian optimization dynamically adjusts the key parameters of the initial task adaptation module, enabling the model to adjust the learning strategy according to the specific task background information, and finally generating an adaptive learning rate adjustment strategy.

[0087] In the embodiment of the present application, first, a meta-learning framework and a reinforcement learning algorithm are used to construct an initial task adaptation module based on a multi-level feedback system. This module optimizes the conversion efficiency between different tasks through the Q-learning algorithm and the policy gradient method to ensure that the model can quickly adapt to new tasks. Then, based on the optimized prediction path, the variational autoencoder is used to analyze the data distribution change trend of different-level feedback signals in the multi-level feedback system, and the key parameters in the initial task adaptation module are dynamically adjusted through Bayesian optimization to generate the final target task adaptation module. Finally, the context awareness mechanism is used to enable the target task adaptation module to adjust the learning strategy according to the background information of different tasks and generate an optimal learning rate adjustment strategy. This process enables the model to flexibly adjust the learning strategy according to the current task background and ensures high-efficiency performance in the face of complex environments.

[0088] The following is a specific example:

[0089] In an application scenario of an intelligent customer service system, users may ask various types of queries, such as product information, order status, and after-sales service. To improve the response speed and accuracy, the task adaptation module first uses a meta-learning framework and a reinforcement learning algorithm to construct an initial module based on a multi-level feedback system and optimizes the conversion efficiency between different tasks through the Q-learning algorithm and the policy gradient method. For example, when a user first asks about product information, the system quickly identifies the user's intention and provides an accurate answer. Then, based on the optimized prediction path, the system uses the variational autoencoder to analyze the data distribution change trend of different-level feedback signals and dynamically adjusts the key parameters through Bayesian optimization to generate the final target task adaptation module. For example, when continuously asking multiple attributes of the same product, the system captures the association between these queries and optimizes the parameter configuration. Finally, the context awareness mechanism is used to adjust the learning strategy according to the background of different tasks and generate an optimal learning rate adjustment strategy. For example, when a user asks about after-sales service, the system quickly adjusts the parameters related to after-sales service to improve the query efficiency. In this way, the system can not only give accurate responses in a short time, but also continuously optimize itself to provide more personalized services.

[0090] To address the problems that the task adaptation module in the existing method is not flexible enough in adjustment and the learning rate adjustment strategy lacks dynamics, in the above-mentioned multiple embodiments, in step 104, by using the context awareness mechanism, the target task adaptation module adjusts the learning strategy according to the background information of different tasks to generate an adaptive learning rate adjustment strategy, and further includes:

[0091] Using the context awareness mechanism, combining graph neural network and attention mechanism, analyzing key features from the background information of different tasks in the multi-level feedback system, where the key features include task type, data distribution characteristics, and historical performance; based on the key features, combining Bayesian optimization and Gaussian process regression, dynamically adjusting the learning rate in the target task adaptation module to generate a preliminary adaptive learning rate adjustment strategy; using a time series prediction model to optimize the preliminary adaptive learning rate adjustment strategy through reinforcement learning to generate an optimized learning rate adjustment strategy; combining genetic algorithm and differential evolution algorithm to search for the optimal learning rate adjustment strategy from the optimized learning rate adjustment strategy; introducing an active learning mechanism to select target data from the optimal learning rate adjustment strategy to generate an adaptive learning rate adjustment strategy. Optionally, the using the context awareness mechanism, combining graph neural network and attention mechanism, analyzing key features from the background information of different tasks in the multi-level feedback system includes: using data mining technology to collect and preprocess the background information of different tasks in the multi-level feedback system to obtain the task background data of the multi-level feedback system; introducing the context awareness mechanism, combining principal component analysis and factor analysis, to perform preliminary screening and dimensionality reduction processing on the task background data to obtain the target task background information; applying graph neural network and complex network analysis to analyze the complex relationships between tasks from the target task background information to construct a task relationship graph; applying the attention mechanism, combining deep reinforcement learning, to highlight the important features in the task relationship graph to obtain the key feature weight assignment result; analyzing the task relationship graph and the key feature weight assignment result to extract key features.

[0092] To address the problems that the analysis of the data distribution change trend in the existing method is not fine enough and the adjustment of key parameters lacks flexibility, in the above-mentioned one or more embodiments, in step 104, based on the optimized prediction path, using a variational autoencoder to analyze the change trend from the data distribution of different-level feedback signals in the multi-level feedback system, and dynamically adjusting the key parameters in the initial task adaptation module through Bayesian optimization to generate a target task adaptation module, and further includes:

[0093] According to the optimized predicted path, using a variational autoencoder and statistical analysis tools, model the data distributions of feedback signals at different levels in the multi-level feedback system, and analyze potential patterns and variation rules from the feedback information to generate a trend of data distribution changes; based on the trend of data distribution changes, through Bayesian optimization, Gaussian process regression, and gradient boosting trees, dynamically adjust the key parameters in the initial task adaptation module to generate a target task adaptation module.

[0094] In this embodiment, the optimized predicted path is used to guide the variational autoencoder and statistical analysis tools to model the data distributions of feedback signals at different levels in the multi-level feedback system. These feedback signals include real-time feedback from user interactions, historical task results, and intermediate representations within the model. Through modeling, the system can capture potential patterns and variation rules and generate an accurate trend of data distribution changes. Subsequently, based on these trends of change, advanced algorithms such as Bayesian optimization, Gaussian process regression, and gradient boosting trees are used to dynamically adjust the key parameters in the initial task adaptation module, and finally a more flexible and efficient target task adaptation module is generated.

[0095] In the embodiment of the present application, first, according to the optimized predicted path, use a variational autoencoder and statistical analysis tools to model the data distributions of feedback signals at different levels in the multi-level feedback system, analyze potential patterns and variation rules to generate a trend of data distribution changes. Then, based on these trends of change, dynamically adjust the key parameters in the initial task adaptation module through Bayesian optimization, Gaussian process regression, and gradient boosting trees to ensure that each parameter is optimally configured according to the latest data distribution changes. This process not only improves the model's adaptability to new tasks but also enhances its stability in complex environments.

[0096] The following is a specific example:

[0097] In an application scenario of an intelligent customer service system, users may pose various types of queries, such as product information, order status, and after-sales service. To improve response speed and accuracy, the system first, according to the optimized predicted path, uses a variational autoencoder and statistical analysis tools to model the data distributions of feedback signals at different levels in the multi-level feedback system. For example, when a user continuously asks multiple attributes of the same product, the system captures the associations between these queries through a variational autoencoder and identifies potential patterns and variation rules through statistical analysis tools to generate a trend of data distribution changes.

[0098] Next, based on these changing trends, the system dynamically adjusts the key parameters in the initial task adaptation module through Bayesian optimization, Gaussian process regression, and gradient boosting trees. For example, when the user starts to frequently ask about after-sales service, the system quickly adjusts the parameters related to after-sales service to optimize the query efficiency. In this way, the system can not only give accurate responses in a short time but also continuously optimize itself as the user's behavior changes, providing more personalized services.

[0099] To address the problems of insufficient depth in multimodal data fusion and inaccurate semantic association analysis in existing methods, in one or more of the above embodiments, in step 102, the use of a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic association and the visual or auditory data related to the text, and determine the multimodal perception mechanism, includes:

[0100] Using a hierarchical graph neural network to define an initial intelligent semantic association network, analyzing the initial semantic association from the internal connection between the text content of the initial intelligent semantic association network and the visual or auditory data associated with the text content; based on the initial semantic association, introducing a cross-modal attention mechanism and combining it with a bidirectional long short-term memory network to perform joint encoding processing on the text and the corresponding visual or auditory data in the initial semantic association to obtain a multimodal information representation; according to the multimodal information representation, constructing a generative adversarial network, analyzing the key data from the adversarial training process between the generator and the discriminator in the generative adversarial network; based on the key data, combining with a variational autoencoder to optimize the initial intelligent semantic association network to obtain a target intelligent semantic association network, and determining the multimodal perception mechanism based on the target intelligent semantic association network.

[0101] In this embodiment, a hierarchical graph neural network (HGNN) is used to define an initial intelligent semantic association network, which can analyze the initial semantic association from the internal connection between the text content and the visual or auditory data associated with it. The initial semantic association refers to the preliminary correspondence relationship between different modal data. Then, a cross-modal attention mechanism is introduced and combined with a bidirectional long short-term memory network (BiLSTM) to perform joint encoding processing on these initial semantic associations to generate a multimodal information representation. This representation not only includes the direct mapping between the text and the visual / audio data but also covers deeper semantic connections. Finally, by constructing a generative adversarial network, key data is extracted from the adversarial training process between the generator and the discriminator, and the initial intelligent semantic association network is optimized by combining with a variational autoencoder to obtain a target intelligent semantic association network, thereby determining the multimodal perception mechanism.

[0102] In the embodiments of the present application, first, a hierarchical graph neural network is used to define an initial intelligent semantic association network, analyze the internal relationship between the text content and the associated visual or auditory data, and generate initial semantic associations. Then, based on these initial semantic associations, a cross-modal attention mechanism combined with BiLSTM is introduced to perform joint encoding processing on the text and the corresponding visual or auditory data to obtain a multi-modal information representation. Next, according to the multi-modal information representation, an adversarial generation network is constructed, and key data is analyzed through the adversarial training process between the generator and the discriminator. Finally, the initial intelligent semantic association network is optimized by combining a variational autoencoder to obtain a target intelligent semantic association network, and finally a multi-modal perception mechanism is determined. This process ensures the depth of multi-modal data fusion and the accuracy of semantic association analysis.

[0103] The following is a specific example:

[0104] In a multimedia content recommendation system, a user may upload a photo or a video clip with a text description. To improve the accuracy and personalization of the recommendation, the system first uses a hierarchical graph neural network to define an initial intelligent semantic association network, analyze the internal relationship between the text description of the photo or video and its visual content, and generate initial semantic associations. For example, when a user uploads a photo labeled with "beach", the system recognizes visual elements such as the beach and the waves and establishes a preliminary association. Then, based on these initial semantic associations, the system introduces a cross-modal attention mechanism combined with BiLSTM to perform joint encoding processing on the text and visual data to generate a multi-modal information representation, which can not only recognize the labels but also understand the specific visual elements. Next, according to the multi-modal information representation, an adversarial generation network is constructed, and key data is extracted through the adversarial training of the generator and the discriminator. Finally, the initial intelligent semantic association network is optimized by combining a variational autoencoder to obtain a target intelligent semantic association network, thereby determining a multi-modal perception mechanism to more accurately understand the content uploaded by the user and provide personalized recommendation services.

[0105] To solve the problems that the incremental update of the context expansion module in the existing methods is not flexible enough and the multi-modal data modeling is not fine enough, in one or more of the above embodiments, in step 105, based on the optimized prediction path and the adaptive learning rate adjustment strategy, an incremental context expansion module is constructed by using the structured representation learning technology based on the graph neural network, including:

[0106] Using the optimized predicted path and the adaptive learning rate adjustment strategy, combining time series analysis and Bayesian optimization, to determine the incremental update rule; based on the incremental update rule, using the structured representation learning technology based on graph neural network, combining with the hierarchical clustering algorithm, to model the multi-modal data in the multi-modal perception mechanism, and obtain the structured representation of multi-modal data; according to the structured representation of multi-modal data, combining the dynamic graph convolutional network and the variational autoencoder, to encode the data obtained from the multi-level feedback system, and input the data into the structured representation of multi-modal data, so as to output the updated context representation; applying the attention mechanism, combining with deep reinforcement learning, to highlight the key features in the updated context representation, and obtain the optimized context representation; combining reinforcement learning and genetic algorithm, to optimize and adjust the parameters in the incremental update process of the task adaptive module, and generate the optimized parameter configuration; based on the optimized context representation and the optimized parameter configuration, to generate the incremental context expansion module.

[0107] In this embodiment, the optimized predicted path and the adaptive learning rate adjustment strategy are used to combine time series analysis and Bayesian optimization to determine the incremental update rule. These rules guide how the system gradually introduces new information without affecting the existing performance. Based on the incremental update rule, the structured representation learning technology based on graph neural network and the hierarchical clustering algorithm are used to model the multi-modal data in the multi-modal perception mechanism, and the structured representation of multi-modal data is obtained. These representations not only capture the associations between different modal data, but also reflect their internal structures. Subsequently, combining the dynamic graph convolutional network and the variational autoencoder, the data obtained from the multi-level feedback system is encoded, and these data are input into the structured representation of multi-modal data, and the updated context representation is output. For further optimization, the attention mechanism is applied in combination with deep reinforcement learning to highlight the key features and generate the optimized context representation. Finally, combining reinforcement learning and genetic algorithm to optimize and adjust the parameters in the incremental update process of the task adaptive module, and generate the optimized parameter configuration, so as to construct the incremental context expansion module.

[0108] In the embodiments of the present application, first, an optimized prediction path and an adaptive learning rate adjustment strategy are utilized, combined with time series analysis and Bayesian optimization, to determine the incremental update rules. Then, based on these rules, a structured representation learning technique based on graph neural networks and a hierarchical clustering algorithm are used to model the multi-modal data in the multi-modal perception mechanism, generating a structured representation of the multi-modal data. Next, a dynamic graph convolutional network and a variational autoencoder are combined to encode the data in the multi-level feedback system, outputting an updated context representation. On this basis, an attention mechanism is applied in combination with deep reinforcement learning to highlight key features, generating an optimized context representation. Finally, the incremental update parameters in the task adaptive module are optimized through reinforcement learning and genetic algorithms, generating an optimized parameter configuration, and finally an incremental context expansion module is constructed. This process ensures the flexibility of the context expansion module and the fineness of multi-modal data modeling.

[0109] The following is a specific example:

[0110] In an intelligent assistant application scenario, a user may continuously ask multiple related questions, such as asking about restaurant information, menu content, and business hours. To improve the response speed and accuracy, the system first uses an optimized prediction path and an adaptive learning rate adjustment strategy, combined with time series analysis and Bayesian optimization, to determine the incremental update rules to guide the gradual introduction of new information without affecting the existing performance. Based on these rules, the system uses graph neural networks and a hierarchical clustering algorithm to model the multi-modal data, generating a structured representation. For example, when the user uploads a photo containing "menu" and asks about nearby restaurants, the system identifies the dish names on the menu and associates them with the queried restaurants. Then, a dynamic graph convolutional network and a variational autoencoder are combined to encode the data in the feedback system, outputting an updated context representation. The attention mechanism and deep reinforcement learning are applied to highlight key features, generating an optimized context representation. Finally, the parameter configuration in the task adaptive module is optimized through reinforcement learning and genetic algorithms to construct an incremental context expansion module, ensuring that the system remains efficient in the face of complex requirements and provides more accurate and personalized answers.

[0111] This application takes into account that in complex multi-modal task processing, traditional static learning rates and fixed parameter configurations are difficult to adapt to the ever-changing task background information. To improve the adaptive ability and generalization performance of the model, a method that can dynamically adjust the learning strategy according to the task background is needed. Especially when facing diverse task types, data distribution characteristics, and historical performance, the model must have the ability to flexibly adjust the learning rate to ensure efficient switching and continuous optimization between different tasks. In addition, existing methods often ignore the impact of task background information on the learning strategy, resulting in poor performance of the model when dealing with new tasks or environmental changes. Therefore, a new alternative solution is proposed, which includes:

[0112] Using the context-aware mechanism, enabling the target task adaptive module to adjust the learning strategy according to the background information of different tasks and generate an adaptive learning rate adjustment strategy, including:

[0113] Using the context-aware mechanism, combining the graph neural network and the attention mechanism, analyzing key features from the background information of different tasks provided by the multi-level feedback system, and the key features are obtained by calculating the attention weights;

[0114] Among them, the calculation formula of the attention weight is as follows:

[0115] ;

[0116] Among them, represents the unnormalized attention score of the th feature or data point, and are the query vector and the key vector respectively, and are the weight matrices, is the bias term, is the adjustment coefficient, is the additional weight matrix, is the context feature vector, is the amplitude coefficient, is the angular frequency, represents the time step, indicating the current time or sequence position, is the phase angle, represents the corresponding attention weight, represents the unnormalized attention score of the th feature or data point, represents the index variable used to traverse all features or data points;

[0117] The key feature is used to adjust the learning strategy in the target task adaptive module to generate an adaptive learning rate adjustment strategy, which is determined by calculating the adaptive learning rate.

[0118] Among them, the adaptive learning rate adjustment formula is as follows:

[0119] ;

[0120] Among them, represents the learning rate at time , represents the initial learning rate, represents the decay coefficient, represents the amplitude, represents the angular frequency, represents the phase angle, and are the progressive growth coefficient and rate respectively, and are additional Gaussian decay coefficients.

[0121] For the formula ;

[0122] The following gives a detailed explanation of each parameter:

[0123] refers to the unnormalized attention score of the th feature or data point. It is used to measure the importance of this feature in the current task.

[0124] and are the query vector and key vector respectively, representing the association between different modality data. Usually obtained by extraction through an encoder.

[0125] and refer to the weight matrices, which are used to perform linear transformations on the query vector and key vector to enhance the expressive power of the model. These matrices are usually learned during the training process.

[0126] refers to the bias term, which is used to adjust the baseline level of the overall score to prevent the output value from being too small or too large. Usually a constant, initially set to 0 or a small value and adjusted during training.

[0127] refers to the adjustment coefficient, which controls the influence degree of context features. Determined through experiments or hyperparameter tuning.

[0128] Refers to an additional weight matrix used for linear transformation of the context feature vector. It is also learned through the training process.

[0129] Refers to the context feature vector, representing background information related to the current task. It can be extracted from a multi-level feedback system.

[0130] Refers to the amplitude coefficient that controls the amplitude of the sine function. It is determined through experiments or hyperparameter tuning.

[0131] Refers to the angular frequency, representing the speed of periodic change. It is set according to the task characteristics. For example, in time series tasks, it may be set according to the time interval.

[0132] Refers to the time step, representing the current time or sequence position. It is directly obtained from the input data.

[0133] Refers to the phase angle, representing the starting position of periodic change. It is set according to the task characteristics. For example, in time series tasks, it may be set according to the initial state.

[0134] Refers to the corresponding attention weight, used to normalize the importance of each feature and ensure that the sum of all weights is 1. This enables the model to allocate computational resources more effectively.

[0135] The design reasons for each sub-item are introduced as follows:

[0136] Query-key similarity term is designed to measure the similarity between different modal data and help the model identify which features are most relevant.

[0137] The design purpose of the bias term (b) is to prevent the output value from being too small or too large and provide additional flexibility.

[0138] Context feature enhancement term is designed to introduce context features and enhance the model's understanding of background information.

[0139] The reason for addition is that as an additional term, it improves the model's sensitivity to complex background information.

[0140] Periodic change term is designed to introduce the time factor, enabling the model to dynamically adjust attention weights, which is particularly important when dealing with time series data or sequence tasks.

[0141] The reason for addition is that as a dynamic adjustment term, it ensures that the model can adapt to the task environment with periodic changes.

[0142] Regarding the formula ;

[0143] The following gives a detailed explanation of each parameter:

[0144] refers to the learning rate at time , which is used to dynamically adjust the learning strategy of the model.

[0145] refers to the initial learning rate, usually determined by experience or experiments, and is the learning rate when the model starts training.

[0146] refers to the decay coefficient, which controls the exponential decay rate of the learning rate over time. It is determined by experiments or hyperparameter tuning.

[0147] refers to the amplitude, which controls the amplitude of the sine wave and introduces periodic changes. It is determined by experiments or hyperparameter tuning.

[0148] refers to the angular frequency, which represents the speed of periodic changes. It is set according to the task characteristics.

[0149] refers to the phase angle, which represents the starting position of periodic changes. It is set according to the task characteristics.

[0150] refers to the progressive growth coefficient, which controls the growth rate of the logarithmic term. It is determined by experiments or hyperparameter tuning.

[0151] refers to the rate, which affects the growth rate of the logarithmic term. It is determined by experiments or hyperparameter tuning.

[0152] refers to the Gaussian decay coefficient, which controls the amplitude of the Gaussian decay term. It is determined by experiments or hyperparameter tuning. Its main role is to mainly affect the initial level and maximum value of the learning rate.

[0153] refers to the Gaussian decay coefficient, which controls the width of the Gaussian decay term. It is determined by experiments or hyperparameter tuning. Its main role is to mainly control the change speed of the learning rate over time.

[0154] The following introduces the design reasons for each sub-item:

[0155] The exponential decay term is designed to simulate the phenomenon that the learning rate gradually decreases over time, which helps the model to be more stable in the later stage of training.

[0156] The sine wave term The design purpose is to introduce periodic changes, increase the learning rate during certain time periods, and help the model explore new solution spaces.

[0157] The reason for addition is as a fluctuation term to provide short-term learning rate adjustment.

[0158] Logarithmic growth term The design purpose is to introduce an asymptotic growth characteristic, make the learning rate rise rapidly at the beginning of training and grow slowly in the later stage, and help the model optimize at different stages.

[0159] The reason for addition is as an asymptotic term to provide mid-term learning rate adjustment.

[0160] Gaussian decay term The design purpose is to introduce Gaussian decay characteristics, make the learning rate larger at the beginning of training and then decline rapidly, and help the model converge quickly in the early stage.

[0161] The overall design aims to construct a flexible and adaptive learning rate adjustment mechanism, enabling the model to dynamically adjust its learning strategy in different task backgrounds. By combining multiple mathematical functions, the formula not only considers the long-term trend over time (such as exponential decay), but also introduces short-term fluctuations (such as sine waves) and mid-term adjustments (such as logarithmic growth), as well as the need for rapid convergence in the early stage (such as Gaussian decay). This multi-dimensional design ensures that the model maintains high performance in the face of complex and changing task environments, ultimately improving the system's response speed and accuracy.

[0162] The following is a specific example:

[0163] In the application scenario of an intelligent customer service system, users may pose various types of queries, such as product information, order status, and after-sales service. To improve the response speed and accuracy, the system needs to dynamically adjust the learning strategy and learning rate according to different task backgrounds.

[0164] Suppose the system is currently processing a query about product information, with the task type being "product information query", the data distribution characteristic being "high-frequency words concentrated in product feature descriptions", and the historical performance being "high frequency of similar queries in the past month".

[0165] Based on the above task background information, the system uses a graph neural network to extract relevant features and calculates the importance of each feature through an attention weight formula. For example:

[0166] and are the query vector and key vector respectively, with values and . and is the weight matrix, with values and .

[0167] is the bias term, with a value of 0.5. is the adjustment coefficient, with a value of 0.1. is the additional weight matrix, with values . is the context feature vector, with values . is the amplitude coefficient, with a value of 0.2. is the angular frequency, with a value of 0.5. is the time step, with a value of 1. is the phase angle, with a value of 0.3.

[0168] Calculation result:

[0169] ;

[0170] After normalization, the attention weight is obtained.

[0171] According to the calculated attention weights, adjust the learning strategy in the target task adaptive module. For example, for key features with high attention weights (such as product feature descriptions), the system will assign higher weights, thus focusing more on the learning of these features.

[0172] According to the adjusted learning strategy, calculate the learning rate using the adaptive learning rate adjustment formula. Assume:

[0173] Initial learning rate ; Decay coefficient ; Amplitude ; Angular frequency ; Phase angle ; Asymptotic growth coefficient ; Rate ; Gaussian decay coefficient .

[0174] Calculation result:

[0175] ;

[0176] As can be seen from the calculation results, the learning rate obtained by the system is 0.01193. Assuming that the range of the learning rate is set to [0, 1] according to the requirements, it indicates that when processing the current task (product information query), the system adopts a relatively poor learning rate in order to adapt to the new task background information more quickly. As time and task complexity change, the learning rate will be gradually adjusted to ensure that the system can maintain efficient conversion and continuous optimization among different tasks. This method not only improves the response speed and accuracy of the system, but also enhances its adaptability in a complex and changing task environment, ultimately improving the user experience and satisfaction.

[0177] Figure 2 The following is a schematic structural diagram of a machine language large model construction system provided by an embodiment of the present application. As Figure 2 shown, the system includes:

[0178] A receiving module 21, configured to receive user input data, define an intelligent semantic association network, and obtain a semantic association based on the intelligent semantic association network and the user input data;

[0179] An enhancement module 22, configured to use a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic association and the visual or auditory data related to the text, and determine a multi-modal perception mechanism;

[0180] A construction module 23, configured to construct a multi-level feedback system based on the multi-modal perception mechanism, and use the multi-level feedback system to process feedback signals from different levels to generate an optimized prediction path;

[0181] An adjustment module 24, configured to construct a task adaptation module based on the multi-level feedback system, using a meta-learning framework and a reinforcement learning algorithm, and adjust key parameters in the task adaptation module based on the task adaptation module and the optimized prediction path to generate an adaptive learning rate adjustment strategy;

[0182] A utilization module 25, configured to construct an incremental context expansion module based on the optimized prediction path and the adaptive learning rate adjustment strategy, using a structured representation learning technology based on a graph neural network, and construct a machine language large model based on the incremental context expansion module, the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptation module.

[0183] Figure 2 The described machine language large model construction system can execute Figure 1A method for constructing a machine language large model described in the illustrated embodiment, the implementation principle and technical effects of which will not be elaborated further. For the machine language large model construction system in the above embodiment, the specific manners in which each module and unit perform operations have been described in detail in the embodiment related to the method, and will not be elaborated herein.

[0184] In a possible design, Figure 2 A machine language large model construction system of the illustrated embodiment can be implemented as a computing device, such as Figure 3 as shown, the computing device may include a storage component 31 and a processing component 32;

[0185] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.

[0186] The processing component 32 is configured to: receive user input data, define an intelligent semantic association network, and based on the intelligent semantic association network and in combination with the user input data, obtain semantic associations; use a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic associations and the visual or auditory data related to the text, and determine a multi-modal perception mechanism; based on the multi-modal perception mechanism, construct a multi-level feedback system, and use the multi-level feedback system to process feedback signals from different levels to generate an optimized prediction path; based on the multi-level feedback system, use a meta-learning framework and a reinforcement learning algorithm to construct a task adaptation module, and based on the task adaptation module and the optimized prediction path, adjust the key parameters in the task adaptation module to generate an adaptive learning rate adjustment strategy; based on the optimized prediction path and the adaptive learning rate adjustment strategy, use a structured representation learning technique based on a graph neural network to construct an incremental context expansion module, and based on the incremental context expansion module, the intelligent semantic association network, the multi-modal perception mechanism, the multi-level feedback system, and the task adaptation module, construct a machine language large model.

[0187] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0188] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0189] Of course, the computing device may also necessarily include other components, such as an input / output interface, a display component, a communication component, and the like.

[0190] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above-mentioned peripheral interface module can be an output device, an input device, etc.

[0191] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0192] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device can refer to a cloud server, and the above-mentioned processing component, storage component, etc. can be basic server resources leased or purchased from a cloud computing platform.

[0193] The embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above-mentioned Figure 1 machine language large model construction method shown in the above embodiment.

[0194] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0195] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0196] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for constructing a large machine language model, characterized in that: include: Receive user input data, and define an intelligent semantic association network, and obtain semantic association based on the intelligent semantic association network combined with the user input data, wherein the user input data includes text, image and audio; Using a cross-modal attention mechanism and a generative adversarial network, the relationship between the text in the semantic association and the visual or auditory data related to the text is enhanced to determine a multimodal perception mechanism; Based on the multimodal perception mechanism, a multi-level feedback system is constructed, and the multi-level feedback system is used to process feedback signals from different levels to generate an optimized prediction path; Based on the multi-level feedback system, a task adaptation module is constructed using a meta-learning framework and a reinforcement learning algorithm, and based on the task adaptation module and the optimized prediction path, key parameters in the task adaptation module are adjusted to generate an adaptive learning rate adjustment strategy; Based on the optimized prediction path and the adaptive learning rate adjustment strategy, an incremental context extension module is constructed using the structured representation learning technology based on graph neural networks. Based on the incremental context extension module, the intelligent semantic association network, the multimodal perception mechanism, the multi-level feedback system and the task adaptation module, a large machine language model is constructed.

2. The method according to claim 1, characterized in that The method of constructing a task adaptation module based on the multi-level feedback system using a meta-learning framework and a reinforcement learning algorithm, and adjusting key parameters in the task adaptation module based on the task adaptation module and the optimized prediction path to generate an adaptive learning rate adjustment strategy includes: Using a meta-learning framework and a reinforcement learning algorithm, an initial task adaptation module is constructed based on the multi-level feedback system, and the conversion efficiency of the initial task adaptation module between different tasks is optimized by a Q-learning algorithm and a policy gradient method; Based on the optimized prediction path, a variational autoencoder is used to analyze the change trend from the data distribution of feedback signals at different levels in the multi-level feedback system, and the key parameters in the initial task adaptation module are dynamically adjusted through Bayesian optimization to generate a target task adaptation module; By using the context-aware mechanism, the target task adaptive module adjusts the learning strategy according to the background information of different tasks to generate an adaptive learning rate adjustment strategy.

3. The method according to claim 2, characterized in that The use of the context-aware mechanism enables the target task adaptive module to adjust the learning strategy according to the background information of different tasks and generate an adaptive learning rate adjustment strategy, including: Utilizing a context-aware mechanism, combined with a graph neural network and an attention mechanism, key features are analyzed from the background information of different tasks in the multi-level feedback system, wherein the key features include task type, data distribution characteristics, and historical performance; Based on the key features, in combination with Bayesian optimization and Gaussian process regression, the learning rate in the target task adaptive module is dynamically adjusted to generate a preliminary adaptive learning rate adjustment strategy; Using a time series prediction model, optimizing the preliminary adaptive learning rate adjustment strategy through reinforcement learning to generate an optimized learning rate adjustment strategy; Combining a genetic algorithm and a differential evolution algorithm, searching for an optimal learning rate adjustment strategy from the optimized learning rate adjustment strategy; An active learning mechanism is introduced to select target data from the optimal learning rate adjustment strategy to generate an adaptive learning rate adjustment strategy.

4. The method according to claim 3, characterized in that The context-aware mechanism is used in combination with the graph neural network and the attention mechanism to analyze key features from the background information of different tasks in the multi-level feedback system, including: Using data mining technology, background information of different tasks in the multi-level feedback system is collected and preprocessed to obtain task background data of the multi-level feedback system; Introducing a context-aware mechanism, combined with principal component analysis and factor analysis, the task background data is preliminarily screened and dimensionally reduced to obtain the target task background information; Applying graph neural networks and complex network analysis, the complex relationships between tasks are analyzed from the background information of the target tasks to construct a task relationship graph; The attention mechanism is applied in combination with deep reinforcement learning to highlight the important features in the task relationship graph and obtain the key feature weight distribution result; The task relationship diagram and the key feature weight distribution result are analyzed to extract key features.

5. The method according to claim 2, characterized in that: Based on the optimized prediction path, the variational autoencoder is used to analyze the change trend from the data distribution of feedback signals at different levels in the multi-level feedback system, and the key parameters in the initial task adaptation module are dynamically adjusted through Bayesian optimization to generate a target task adaptation module, including: According to the optimized prediction path, using variational autoencoders and statistical analysis tools, modeling the data distribution of feedback signals at different levels in the multi-level feedback system, and analyzing potential patterns and change rules from the feedback information to generate data distribution change trends; Based on the data distribution change trend, the key parameters in the initial task adaptive module are dynamically adjusted through Bayesian optimization, Gaussian process regression and gradient boosting tree to generate a target task adaptive module.

6. The method according to claim 1, characterized in that The method utilizes a cross-modal attention mechanism and a generative adversarial network to enhance the relationship between the text in the semantic association and the visual or auditory data related to the text, and determines a multimodal perception mechanism, including: Using a hierarchical graph neural network, an initial intelligent semantic association network is defined, and an initial semantic association is analyzed from the intrinsic connection between the text content of the initial intelligent semantic association network and the visual or auditory data associated with the text content; Based on the initial semantic association, a cross-modal attention mechanism is introduced, and a bidirectional long short-term memory network is combined to jointly encode the text in the initial semantic association and the corresponding visual or auditory data to obtain a multimodal information representation; Constructing a generative adversarial network based on the multimodal information representation, and analyzing key data from an adversarial training process between a generator and a discriminator in the generative adversarial network; Based on the key data and in combination with a variational autoencoder, the initial intelligent semantic association network is optimized to obtain a target intelligent semantic association network, and based on the target intelligent semantic association network, a multimodal perception mechanism is determined.

7. The method according to claim 1, characterized in that The method of constructing an incremental context expansion module based on the optimized prediction path and the adaptive learning rate adjustment strategy by using a structured representation learning technology based on a graph neural network includes: Determine an incremental update rule by using the optimized prediction path and the adaptive learning rate adjustment strategy in combination with time series analysis and Bayesian optimization; Based on the incremental update rule, the multimodal data in the multimodal perception mechanism is modeled and processed by using a structured representation learning technology based on a graph neural network in combination with a hierarchical clustering algorithm to obtain a structured representation of the multimodal data; According to the multimodal data structured representation, in combination with a dynamic graph convolutional network and a variational autoencoder, the data obtained from the multi-level feedback system is encoded, and the data is input into the multimodal data structured representation to output an updated context representation; Applying an attention mechanism in combination with deep reinforcement learning to highlight key features in the updated context representation to obtain an optimized context representation; Combining reinforcement learning and genetic algorithms, optimizing and adjusting the parameters in the incremental update process of the task adaptive module to generate an optimized parameter configuration; An incremental context extension module is generated based on the optimized context representation and the optimized parameter configuration.

8. A machine language large model construction system, characterized in that: include: A receiving module, used to receive user input data, and define an intelligent semantic association network, and obtain semantic association based on the intelligent semantic association network combined with the user input data, wherein the user input data includes text, image and audio; An enhancement module, for enhancing the relationship between the text in the semantic association and the visual or auditory data related to the text by using a cross-modal attention mechanism and a generative adversarial network, and determining a multimodal perception mechanism; A construction module, used to construct a multi-level feedback system based on the multimodal perception mechanism, and use the multi-level feedback system to process feedback signals from different levels to generate an optimized prediction path; An adjustment module is used to construct a task adaptation module based on the multi-level feedback system using a meta-learning framework and a reinforcement learning algorithm, and to adjust key parameters in the task adaptation module based on the task adaptation module and the optimized prediction path to generate an adaptive learning rate adjustment strategy; A module is used to construct an incremental context extension module based on the optimized prediction path and the adaptive learning rate adjustment strategy using a structured representation learning technology based on a graph neural network, and a large machine language model is constructed based on the incremental context extension module, the intelligent semantic association network, the multimodal perception mechanism, the multi-level feedback system and the task adaptation module.

9. A computing device, characterized in that It comprises a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a machine language large model construction method as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, a method for constructing a large machine language model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Methods and apparatus for autonomous robotic control

    CA2941250A1

  • Semantic association modeling and self-adaptive simulation method based on multilayer correction

    CN119203607A