Multi-modal task text processing optimization method based on large language model
By constructing a task-based text processing optimization system, coordinating the information interaction between large language models and non-language artificial intelligence models, and performing text semantic optimization and generation result analysis, the system solves the controllability and consistency issues of the generation process of large language models in multimodal tasks, thereby improving the quality and comprehension of text processing.
Patent Information
- Application Number
- CN202511146302.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-12-30
AI Technical Summary
Existing large language models lack quality feedback and early warning mechanisms for the generation process in multimodal tasks, making it difficult to meet customized needs. They are highly versatile but have weak controllability and are difficult to dynamically optimize output content according to actual business rules.
By constructing a task text processing optimization system, including a large language model access end, a multimodal task text data management module, a dynamic management optimization end, and a processing optimization and feedback end, the system coordinates the information interaction between the large language model and non-language artificial intelligence models, performs text semantic optimization, modal alignment, and generation result analysis, and provides visual feedback and early warning mechanisms.
It significantly improves the quality of text processing in multimodal tasks, reduces multimodal alignment error, enhances the contextual consistency and semantic integrity of generated text, and strengthens the understanding and controllability of large language models in multimodal scenarios.
Smart Images

Figure CN121234902A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer text processing optimization, and particularly to a multi-modal task text processing optimization method based on a large language model. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, large language models (LLM) have shown excellent text generation and understanding capabilities in the field of natural language processing (NLP). In recent years, large language models represented by GPT and BERT have been gradually introduced into multi-modal task processing scenarios, such as image-text question answering, speech description generation, image content summarization, and video event recognition. These tasks usually require the model to process multiple modal information inputs (such as text, images, speech, etc.) and output highly relevant comprehensive descriptions or responses in natural language.
[0003] Chinese Patent Publication No. CN 120104781 A discloses a text processing method and system based on a large language model, relating to the field of computer technology, which includes collecting text data, generating variant samples through data enhancement technology, obtaining a pre-training data set, introducing a knowledge graph, and outputting a trained large language model; using an adaptive task selector, an incremental learning framework, and a meta-learning algorithm to optimize hyperparameters; considering non-text data, using a multi-modal model to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expression methods; based on the updated large language model, establishing a confidence evaluation method, and outputting high-confidence text processing results; according to the high-confidence text processing results, constructing a personalized user portrait and optimizing recommended content; the present application effectively improves the performance of the large language model in handling complex real-world problems by introducing a multi-modal model and a cross-cultural cognitive framework.
[0004] However, the above-mentioned scheme still has the following problems: most current multi-modal models mainly focus on end-to-end generation capabilities, lack quality feedback and early warning means for the generation process, and different scenarios have clear requirements for the style, structure, and word usage specifications of text output. The existing methods are difficult to dynamically optimize the output content according to actual business rules, resulting in strong model versatility but weak controllability, making it difficult to meet customized needs. Therefore, the present application needs to design a multi-modal task text processing optimization method based on a large language model to solve the above-mentioned problems. SUMMARY The purpose of the present application is to provide a multi-modal task text processing optimization method based on a large language model to solve the above-mentioned problems.
[0005] To solve the above problems, the present application provides a technical solution: A multi-modal task text processing optimization method based on a large language model, comprising the following specific steps: S1, connect and call a large language model deployed locally or remotely, encapsulate the text input into a request structure conforming to the input format of the large language model, access a non-language artificial intelligence model that cooperates with the large language model, execute core natural language processing tasks based on the optimized input text, coordinate information interaction, task allocation and output fusion between the large language model and other models; S2, optimize the semantic integrity, logicality and readability of the generated text, improve the understanding effect of the large language model, supplement missing information based on the context to improve semantic coherence, preprocess and feature extract the original multi-modal data, automatically identify the modal type in the input data, and use a pre-training encoder to extract vector representation of each modal; S3, manage information selection, cropping and completion of long text context, adapt to the length limit of the large language model, realize accurate semantic alignment between text and modal, and set and control parameters such as artificial pre-defined rules, templates, prompt information and priority logic for various processing scenarios in multi-modal tasks; S4, multi-level quality analysis and content optimization of the generated results in the multi-modal task execution process, consistency comparison of the text results generated corresponding to each modal input content based on semantic similarity algorithm or language model comparison analysis mechanism, supporting users to score and feedback the results through a visual imaging platform, and displaying data in table form.
[0006] As a preferred embodiment of the application, the other models in step S1 include GPT series, GLM, Baichuan, ERNIE and other open source or commercial large language models, which are connected through API, SDK or local inference framework.
[0007] As a preferred embodiment of the application, the automatic identification of the modal type in the input data in step S2 includes images, audio, video, etc., and the pre-training encoder in step S2 includes CLIP, ViT, CNN, WaveNet.
[0008] As a preferred embodiment of the application, the step S3 includes introducing artificially defined classification labels, keywords, forbidden words or semantic rules for filtering, modifying or enhancing processing before, during and after the large language model generates text, to improve system stability and controllability.
[0009] As a preferred embodiment of the application, the step S4 includes image, voice and text.
[0010] As a preferred embodiment of the present application, a task text processing optimization system needs to be constructed before the step S1 is performed, the task text processing optimization system comprising a large language model access end, a multi-modal task text data management module, a dynamic management optimization end and a processing optimization and feedback end, the output ends of the large language model access end and the multi-modal task text data management module being in communication connection with the input end of the dynamic management optimization end, and the output end of the dynamic management optimization end being in communication connection with the input end of the processing optimization and feedback end.
[0011] As a preferred embodiment of the present application, the large language model access end comprises a large language model access unit, other model access units and a model synchronization optimization unit, the output ends of the large language model access unit and the other model access units being in communication connection with the input end of the model synchronization optimization unit. The large language model access unit is used to connect and call a large language model deployed locally or remotely to support natural language processing tasks and encapsulate text input into a request structure conforming to the input format of the large language model. The other model access units are used to access non-language artificial intelligence models that work in cooperation with the large language model to assist or enhance multi-modal task processing capabilities. The non-language artificial intelligence models comprise image recognition models, voice recognition models, video summary models, OCR models and the like, obtain structured semantic information of non-text modalities, share structured results generated by other models to the input process of the large language model as part of prompt embedding, or use the structured results to optimize the generated text output. The model synchronization optimization unit is used to execute core natural language processing tasks based on the optimized input text, coordinate information interaction, task allocation and output fusion between the large language model and other models, improve cross-model collaboration effect, identify semantic conflicts existing in multiple model outputs, and correct or annotate the final text generation result.
[0012] As a preferred embodiment of the present application, the multi-modal task text data management module comprises a text semantic optimization unit, a multi-modal input control unit and a data acquisition positioning unit, the output end of the text semantic optimization unit being in communication connection with the input end of the multi-modal input control unit. The text semantic optimization unit is used to optimize semantic integrity, logicality and readability of generated text and improve understanding effect of the large language model. The text semantic optimization unit is also used to supplement missing information based on context to improve semantic coherence, call external knowledge base or introduce related domain knowledge by using a preset ontology. The multimodal input control unit is used to preprocess and extract features from the raw multimodal data to prepare for subsequent text generation and automatically identify the modality type in the input data. The data acquisition and positioning unit is used to perform data source positioning processing when acquiring task text data, and can locate the data source immediately when data is at risk or abnormal.
[0013] In a preferred embodiment of the present invention, the dynamic management optimization terminal includes a dynamic management optimization unit, a modal alignment control unit, and a manual preset management unit, wherein the dynamic management optimization unit and the modal alignment control unit are bidirectionally connected. The dynamic management and optimization unit is used to manage the selection, trimming and completion of information in the context of long texts, adapting to the length limitations of large language models. Score and sort contextual information to identify key content segments; Delete low-weight or irrelevant paragraphs while preserving the semantic thread; Intelligent completion of incomplete context, such as predicting causal relationships and temporal order; The modality alignment control unit is used to achieve precise semantic alignment between text and modality, thereby improving cross-modal expression consistency. Detect references in the text that are related to the modal content; The references are corrected based on modal feature semantics to ensure accurate semantic pointing. Set the priority or confidence level of each modality in task processing; The manual preset management unit is used to set and control parameters such as manually predefined rules, templates, prompts, and priority logic for various processing scenarios in multimodal tasks, so as to guide, intervene in and limit the large language model and its auxiliary modules. This includes introducing manually defined classification labels, keywords, banned words, or semantic rules for filtering, modification, or enhancement processing before, during, and after text generation by the large language model, thereby improving system stability and controllability.
[0014] In a preferred embodiment of the present invention, the processing optimization and feedback terminal includes a processing optimization unit, a visualization feedback unit, and a field early warning unit, wherein the output terminals of the processing optimization unit and the field early warning unit are both communicatively connected to the input terminal of the visualization feedback unit. The processing optimization unit is used to perform multi-level quality analysis and content optimization on the generated results during the execution of multimodal tasks; The processing optimization unit is also used to perform consistency comparison on the text results generated corresponding to each modal input content based on semantic similarity algorithm or language model comparison analysis mechanism, and to discover and automatically correct content that is contradictory, missing or deviates from the topic. The visualization feedback unit is used to provide a visualization imaging platform and display data in a tabular format. The visualization feedback unit is also used to support users in scoring and providing feedback on results, and to improve prompting strategies and parameter adjustments; The on-site early warning unit is used to monitor and provide early warnings in real time for potential anomalies, model error outputs, modal conflicts, or deviations from key business indicators during the execution of multimodal task processing, so as to ensure the security, reliability, and business adaptability of the output results. The on-site early warning unit is also used to automatically issue a conflict prompt and pause automatic generation when there are obvious conflicts or inability to merge multimodal inputs, in order to prevent the generation of misleading content. The on-site early warning unit also feeds back the early warning information to the visualization feedback unit in the form of pop-ups, voice prompts, and visual icons to ensure closed-loop management of the processing flow.
[0015] The beneficial effects of this invention are as follows: By setting up a large language model access terminal, a multimodal task text data management module, a dynamic management optimization terminal, and a processing optimization and feedback terminal, this invention constructs a complete task text processing optimization system. In actual use, it calls the large language model and other models, encapsulates multimodal input requests and coordinates interactions, optimizes text semantic coherence, automatically identifies modalities and extracts feature vectors, uses a pre-trained encoder to extract vector representations of each modality, achieves cross-modal semantic alignment by adapting to long text length limitations, sets and controls parameters such as manually predefined rules, templates, prompts, and priority logic, analyzes the consistency of generated results by analyzing the input content of each modality, and provides a visual user feedback interface. This significantly improves the text processing quality in multimodal tasks, reduces multimodal alignment errors, improves the contextual consistency and semantic integrity of generated text, enhances the understanding and generalization ability of the large language model in multimodal scenarios, and manages, visualizes, and stores text processing optimization data and corresponding analysis results. This also helps to achieve energy management through IoT cloud control and improves the intelligence level of text processing optimization management. Attached Figure Description For ease of explanation, the present invention will be described in detail below with reference to specific embodiments and accompanying drawings.
[0016] Figure 1 This is an overall flowchart of a multimodal task text processing optimization method based on a large language model according to the present invention. Detailed Implementation like Figure 1 As shown, the specific implementation adopts the following technical solution: An optimization method for multimodal task text processing based on a large language model includes the following specific steps: S1. Connect to and call the large language model deployed locally or remotely, encapsulate the text input into a request structure that conforms to the input format of the large language model, access the non-language artificial intelligence model that works in collaboration with the large language model, perform the core natural language processing tasks based on the optimized input text, and coordinate the information interaction, task allocation and output fusion between the large language model and other models. Other models include open-source or commercial large language models such as GPT series, GLM, Baichuan, and ERNIE, which can be connected via API, SDK or local inference framework; S2. Optimize the semantic integrity, logic and readability of the generated text, improve the understanding effect of the large language model, supplement missing information based on context to improve semantic coherence, preprocess and extract features from the original multimodal data, automatically identify the modality type in the input data, and use the pre-trained encoder to extract the vector representation of each modality. The modality type in the input data is automatically identified, including images, audio, video, etc. The pre-trained encoder in step S2 includes CLIP, ViT, CNN, and WaveNet. S3 manages the selection, trimming, and completion of information in the context of long texts, adapts to the length limitations of large language models, achieves precise semantic alignment between text and modality, and sets and controls parameters such as manually predefined rules, templates, prompts, and priority logic for various processing scenarios in multimodal tasks. When performing various predefined data, including the introduction of manually defined classification labels, keywords, banned words or semantic rules, it is used for filtering, modification or enhancement processing before, during and after the large language model generates text, thereby improving the system's stability and controllability. S4. Perform multi-level quality analysis and content optimization on the generated results during the execution of multimodal tasks. Based on semantic similarity algorithms or language model comparison analysis mechanisms, perform consistency comparison on the text results generated corresponding to each modal input content. Support users to score and provide feedback on the results by providing a visualization imaging platform. Display the data in a table format. The input content for each modality includes images, voice, and text.
[0017] Before performing step S1, a task text processing optimization system needs to be constructed. The task text processing optimization system includes a large language model access terminal, a multimodal task text data management module, a dynamic management optimization terminal, and a processing optimization and feedback terminal. The output terminals of the large language model access terminal and the multimodal task text data management module are both connected to the input terminal of the dynamic management optimization terminal, and the output terminal of the dynamic management optimization terminal is connected to the input terminal of the processing optimization and feedback terminal.
[0018] The large language model access terminal includes a large language model access unit, other model access units, and a model synchronization optimization unit. The outputs of the large language model access unit and other model access units are communicatively connected to the input of the model synchronization optimization unit. The large language model access unit connects to and invokes locally or remotely deployed large language models to support natural language processing tasks, encapsulating text input (including prompts, contextual information, modal descriptions, etc.) into a request structure conforming to the large language model's input format. The other model access units access non-language AI models that work collaboratively with the large language model to assist or enhance multimodal task processing capabilities. These non-language AI models include image recognition models (such as CLIP, Re...). The system utilizes various models, including sNet, ViT, speech recognition models (such as Wav2Vec and Whisper), video summarization models, and OCR models, to acquire structured semantic information from non-textual modalities. It shares the structured results generated by other models with the large language model's input flow as part of the prompt embedding or to optimize the generated text output. The model synchronization optimization unit performs core natural language processing tasks based on the optimized input text, coordinating information interaction, task allocation, and output fusion between the large language model and other models to improve cross-model collaboration. It also identifies semantic conflicts in the outputs of multiple models (e.g., images without "cat" but text mentioning "cat") and corrects or annotates the final text generation results.
[0019] The multimodal task text data management module includes a text semantic optimization unit, a multimodal input control unit, and a data acquisition and positioning unit. The output of the text semantic optimization unit is communicatively connected to the input of the multimodal input control unit. The text semantic optimization unit is used to optimize the semantic completeness, logic, and readability of the generated text, thereby improving the understanding effect of the large language model. The text semantic optimization unit is also used to supplement missing information (such as names, locations, and times) based on context to improve semantic coherence, and to call external knowledge bases or preset ontology to introduce relevant domain knowledge. The multimodal input control unit is used to preprocess and extract features from the original multimodal data to prepare for subsequent text generation, and to automatically identify the modality type in the input data. The data acquisition and positioning unit is used to perform data source positioning processing when acquiring task text data, and can locate the data source immediately when data risks or anomalies occur.
[0020] The dynamic management and optimization terminal includes a dynamic management and optimization unit, a modal alignment control unit, and a manually preset management unit. The dynamic management and optimization unit and the modal alignment control unit are bidirectionally connected. The dynamic management and optimization unit is used to manage the selection, trimming, and completion of information in long text contexts to adapt to the length limitations of large language models; it scores and sorts context information to identify key content fragments; it deletes low-weight or irrelevant paragraphs to maintain the semantic thread; and it intelligently completes incomplete contexts, such as predicting causal relationships and temporal order. The modal alignment control unit is used to achieve accurate semantic alignment between text and modality, improving cross-modal expression consistency. The system detects references related to modal content in the text; corrects references based on modal semantic features to ensure accurate semantic pointing; sets the priority or confidence level of each modality in task processing; the manually preset management unit is used to set and control parameters such as manually predefined rules, templates, prompts, and priority logic for various processing scenarios in multimodal tasks, so as to guide, intervene, and limit the large language model and its auxiliary modules; including the introduction of manually defined classification tags, keywords, banned words, or semantic rules for filtering, modification, or enhancement processing before, during, and after text generation by the large language model, improving system stability and controllability.
[0021] The processing optimization and feedback unit includes a processing optimization unit, a visualization feedback unit, and a field early warning unit. The outputs of the processing optimization unit and the field early warning unit are communicatively connected to the input of the visualization feedback unit. The processing optimization unit performs multi-level quality analysis and content optimization on the generated results during the execution of the multimodal task. It also performs consistency comparison of the generated text results corresponding to the input content of each modality based on semantic similarity algorithms or language model comparison analysis mechanisms, identifying and automatically correcting information contradictions, omissions, or deviations from the topic. The visualization feedback unit provides a visualization imaging platform, displaying data in tabular form. It also supports... The system provides user-scored and feedback results to improve prompting strategies and parameter adjustments. The on-site early warning unit monitors and provides real-time warnings for potential anomalies, model error outputs, modal conflicts, or deviations from key business indicators during multimodal task processing, ensuring the security, reliability, and business adaptability of the output results. The on-site early warning unit also automatically issues conflict warnings and pauses automatic generation when there are significant conflicts or incompatibility between multimodal inputs (e.g., content in an image and audio or text content), preventing the generation of misleading content. Furthermore, the on-site early warning unit feeds back warning information to the visualization feedback unit via pop-ups, voice prompts, and visual icons, ensuring closed-loop management of the processing flow.
[0022] Example When a large language model is connected to a text processing optimization system management: S1. After staff inspected the on-site and remote power equipment one by one and confirmed that they were all in normal working order, the task text processing optimization system was activated. The task text processing optimization system controlled the large language model access unit to connect to and call the large language model deployed locally or remotely to support natural language processing tasks. It encapsulated prompt words, context information, and modal descriptions into a request structure that conformed to the input format of the large language model. The task text processing optimization system controlled other model access units to access non-language artificial intelligence models that worked in collaboration with the large language model to assist or enhance multimodal task processing capabilities. It shared the structured results generated by other models into the input process of the large language model as part of the prompt embedding or to optimize the generated text output. The task text processing optimization system controlled the model synchronization optimization unit to execute the core natural language processing task based on the optimized input text, coordinate the information interaction, task allocation, and output fusion between the large language model and other models, improve the cross-model collaboration effect, identify semantic conflicts in the output of multiple models, and correct or annotate the final text generation result. S2. The task text processing optimization system controls the text semantic optimization unit to optimize the semantic integrity, logic, and readability of the generated text, improve the understanding effect of the large language model, supplement missing names and location information based on context to improve semantic coherence, and call external knowledge bases or preset ontology to introduce relevant domain knowledge. The task text processing optimization system controls the multimodal input control unit to preprocess and extract features from the original multimodal data to prepare for subsequent text generation, automatically identify the modality type in the input data, and control the data acquisition and positioning unit to perform data source positioning processing when acquiring task text data. When data risks or anomalies occur, the data source can be located immediately. S3, the task text processing optimization system controls the dynamic management optimization unit to manage the selection, trimming and completion of information in the context of long texts, adapting to the length limitations of large language models. The task text processing optimization system controls the modal alignment control unit to achieve precise semantic alignment between text and modality, improving cross-modal expression consistency. The task text processing optimization system controls the manual preset management unit to set and adjust parameters such as manually predefined rules, templates, prompts, and priority logic for various processing scenarios in multimodal tasks, so as to guide, intervene and limit the large language model and its auxiliary modules. S4. The task text processing optimization system control and optimization unit performs multi-level quality analysis and content optimization on the generated results during the execution of multimodal tasks. Based on semantic similarity algorithms or language model comparison analysis mechanisms, it performs consistency comparison on the text results generated corresponding to each modal input content, discovers information contradictions, omissions, or deviations from the topic, and automatically corrects them. The task text processing optimization system control visualization feedback unit provides a visualization imaging platform, displays data in tabular form, supports user scoring and feedback on results, and improves prompt strategies and parameter adjustments. During the execution of multimodal task processing, the task text processing optimization system control on-site early warning unit monitors and provides real-time early warning prompts for potential anomalies, model error outputs, modal conflicts, or deviations from key business indicators, ensuring the safety, reliability, and business adaptability of the output results. When there are obvious conflicts or inability to merge multimodal inputs, it automatically issues a conflict prompt and pauses automatic generation to prevent the generation of misleading content. The early warning information is fed back to the visualization feedback unit in the form of pop-ups, voice prompts, and visual icons to ensure closed-loop management of the processing flow.
[0023] Specifically, in practical applications, multiple multimodal task text data management modules are used in conjunction with the large language model access end, dynamic management and optimization end, and processing optimization and feedback end. These multiple multimodal task text data management modules are located in different geographical locations. This invention constructs a complete task text processing and optimization system by setting up a large language model access end, multimodal task text data management modules, dynamic management and optimization end, and processing optimization and feedback end. In actual use, it calls the large language model and other models, encapsulates multimodal input requests and coordinates interactions, optimizes text semantic coherence, automatically identifies modalities and extracts feature vectors, and uses a pre-trained encoder to extract each modality. The vector representation, by adapting to the length limit of long texts, achieves cross-modal semantic alignment, allows for the setting and control of parameters such as manually predefined rules, templates, prompts, and priority logic, analyzes the consistency of generated results by analyzing the input content of each modality, provides a visual user feedback interface, significantly improves the text processing quality in multimodal tasks, reduces multimodal alignment errors, improves the contextual consistency and semantic integrity of generated text, enhances the understanding and generalization ability of large language models in multimodal scenarios, and manages, visualizes, and stores text processing optimization data and corresponding analysis results. This facilitates energy management through IoT cloud control and improves the level of intelligence in text processing optimization management.
[0024] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0025] The modules serving as the large language model access point, multimodal task text data management point, dynamic management optimization point, and processing optimization and feedback point may or may not be physically separate. The components displayed as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of these units can be selected to achieve the purpose of this embodiment according to actual needs.
[0026] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0027] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program instructions, such as USB flash drives, portable hard drives, read-only storage servers, random access storage servers, magnetic disks, or optical disks.
[0028] Furthermore, it should be noted that the combination of the various technical features in this case is not limited to the combination methods described in the claims of this case or the combination methods described in the specific embodiments. All technical features described in this case can be freely combined or combined in any way, unless they contradict each other.
[0029] It should be noted that the above examples are merely specific embodiments of the present invention, and the present invention is obviously not limited to the above embodiments, with many similar variations. All modifications that can be directly derived or conceived by those skilled in the art from the content disclosed in this invention should fall within the protection scope of this invention.
[0030] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A large language model-based multi-modal task text processing optimization method, characterized in that, Comprise the following specific steps: S1, model cooperative control: calling large language model and other models, encapsulating multi-modal input request and coordinating interaction; S2, semantic enhancement processing: optimizing text semantic coherence, automatically identifying modalities and extracting feature vectors, using pre-trained encoder to extract vector representation of each modality; S3, context management: adapting to long text length limit, realizing cross-modal semantic alignment, setting and regulating parameters such as artificial pre-defined rules, templates, prompt information and priority logic; S4, quality optimization control: analyzing the consistency of the generated results of each modality input content, providing visual user feedback interface.
2. The method of claim 1, wherein the method is performed by a computer system. The other models in step S1 include GPT series, GLM, Baichuan, ERNIE and other open source or commercial large language models, which are connected by API, SDK or local inference framework.
3. The multimodal task text processing optimization method based on a large language model according to claim 1, characterized in that: The automatic identification of modal types in the input data in step S2 includes image, audio, video and the like, and the pre-trained encoder in step S2 includes CLIP, ViT, CNN and WaveNet.
4. The method of claim 1, wherein the method is performed by a computer system. In step S3, the introduction of artificial definition of classification labels, keywords, forbidden words or semantic rules is used for filtering, modifying or enhancing processing before, during and after the generation of text by large language model, so as to improve the stability and controllability of the system.
5. The method of claim 1, wherein: The step S4 includes image, speech and text.
6. The method of claim 5, wherein the method further comprises: Before performing the step S1, a task text processing optimization system needs to be constructed, which comprises a large language model access end, a multi-modal task text data management module, a dynamic management optimization end and a processing optimization and feedback end, the output ends of the large language model access end and the multi-modal task text data management module are in communication connection with the input end of the dynamic management optimization end, and the output end of the dynamic management optimization end is in communication connection with the input end of the processing optimization and feedback end.
7. The method of claim 6, wherein the method further comprises: The large language model access end comprises a large language model access unit, an other model access unit and a model synchronization optimization unit, the output ends of the large language model access unit and the other model access unit are in communication connection with the input end of the model synchronization optimization unit; The large language model access unit is used to connect and call the locally or remotely deployed large language model, and encapsulate the text input into a request structure conforming to the input format of the large language model; The other model access unit is used to access non-language artificial intelligence models that work with the large language model; The model synchronization optimization unit is used to execute core natural language processing tasks based on the optimized input text, coordinate information interaction, task allocation and output fusion between the large language model and other models.
8. The method of claim 7, wherein the method further comprises: The multi-modal task text data management module comprises a text semantic optimization unit, a multi-modal input control unit and a data acquisition positioning unit, the output end of the text semantic optimization unit is in communication connection with the input end of the multi-modal input control unit; The text semantic optimization unit is used to optimize the semantic integrity, logicality and readability of the generated text, and improve the understanding effect of the large language model; The multi-modal input control unit is used for pre-processing and feature extraction of original multi-modal data, and automatically identifies the modal type in the input data; The data acquisition positioning unit is used for data source positioning processing when task text data is acquired.
9. The method of claim 6, wherein the method further comprises: determining a text length of the text input; and determining a text length of the text output. The dynamic management optimization end includes a dynamic management optimization unit, a modal alignment control unit, and an artificial preset management unit, and the dynamic management optimization unit and the modal alignment control unit are bidirectionally communicated and connected; The dynamic management optimization unit is used for managing information selection, cropping, and completion of long text context; The modal alignment control unit is used for realizing accurate semantic alignment between text and modal; The artificial preset management unit is used for setting and regulating parameters such as artificial predefined rules, templates, prompt information, and priority logic for various processing scenarios in multi-modal tasks.
10. The multimodal task text processing optimization method based on a large language model according to claim 6, characterized in that: The processing optimization and feedback end includes a processing optimization unit, a visual feedback unit, and a field early warning unit, and the output ends of the processing optimization unit and the field early warning unit are communicated and connected with the input end of the visual feedback unit; The processing optimization unit is used for multi-level quality analysis and content tuning of the generated results in the multi-modal task execution process; The visual feedback unit is used for providing a visual imaging platform and displaying data in a table manner; The field early warning unit is used for real-time monitoring and early warning prompt of potential abnormalities, model error output, modal conflict, or business key indicator deviation in the multi-modal task processing execution process.
Citation Information
Patent Citations
Text processing method and system based on large language model
CN120104781A