Text processing method and system based on large language model
By introducing multimodal models and cross-cultural cognitive frameworks into large language models, combining data augmentation and adaptive task selectors, the shortcomings of large language models in multimodal and cross-cultural scenarios are solved, and more efficient text processing and more accurate user experience are achieved.
Patent Information
- Application Number
- CN202510099874.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing text processing methods based on large language models are insufficient in processing multimodal data and cross-cultural scenarios, and it is difficult to capture multimodal information such as images and tables and understand expression habits of different cultures, and the adaptive task selection mechanism is insufficient in flexibility and efficiency.
By introducing multimodal models and cross-cultural cognitive frameworks, text data is collected and enhanced, knowledge graphs and adaptive task selectors are used to optimize large language models, and hyperparameter optimization is combined with incremental learning and meta-learning algorithms to capture non-text data cues and understand cross-cultural expression.
It improves the performance of large language models when dealing with complex real-world problems, enhances the ability to understand context and cross-cultural expression, improves the applicability and accuracy of application scenarios, and optimizes recommended content through personalized user portraits, significantly improving the user experience.
Smart Images

Figure CN120104781A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a text processing method and system based on a large language model. Background Art
[0002] In recent years, with the rapid development of natural language processing (NLP) technology, text processing methods based on large language models have gradually become a research hotspot. In the existing technology, pre-trained language models such as BERT and RoBERTa have achieved remarkable results on general corpora and are widely used in various downstream tasks. However, the application of these models in specific fields or cross-cultural scenarios still faces challenges. On the one hand, traditional methods have limited processing of non-text data and it is difficult to capture the rich contextual clues brought by multimodal information such as images and tables; on the other hand, the lack of understanding of the expression habits and metaphors of different cultures leads to poor adaptability of the model in global application scenarios. In addition, the existing adaptive task selection mechanism usually relies on static configuration and cannot dynamically respond to changes in the characteristics of the data set, which limits its flexibility and efficiency. Summary of the invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides a text processing method based on a large language model to solve the problems of insufficient multimodal data processing and cross-cultural understanding.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] In a first aspect, the present invention provides a text processing method based on a large language model, which includes collecting text data, generating variant samples through data enhancement technology, obtaining a pre-training data set, introducing a knowledge graph, and outputting a trained large language model;
[0007] Adopting adaptive task selector, the pre-training dataset is automatically matched to the downstream task type. At the same time, the incremental learning framework and meta-learning algorithm are used to optimize hyperparameters.
[0008] Consider non-textual data, use multimodal models to capture different types of contextual clues, and use cross-cultural cognitive frameworks to understand expressions;
[0009] Based on the updated large language model, a confidence assessment method is established and high-confidence text processing results are output;
[0010] Based on high-confidence text processing results, build personalized user portraits and optimize recommended content.
[0011] As a preferred solution of the text processing method based on the large language model of the present invention, wherein: the text data includes authoritative documents, professional forums and academic journals;
[0012] The non-text data includes images and tables.
[0013] As a preferred solution of the text processing method based on the large language model described in the present invention, the text data is collected, and variant samples are generated by data enhancement technology to obtain a pre-training data set, introduce a knowledge graph, and output the trained large language model, including the following steps:
[0014] Determine the target area and identify the main information sources within the target area;
[0015] Use crawler technology to access selected primary information sources and download text data;
[0016] Use synonym replacement, sentence reorganization and context expansion to obtain variant samples;
[0017] Integrate text data and variant samples to generate pre-training datasets;
[0018] Analyze the unique expression habits, metaphors and slang of various cultures to form a cultural feature label library, and use bilingual comparison to build a parallel corpus;
[0019] Apply the BERT-based NER model to automatically identify entities in text and label entity types;
[0020] Through dependency parsing, the grammatical structure of the sentence is analyzed to find the potential relationship between entities;
[0021] The TransE algorithm is used to optimize triple representation and map entities and relations into low-dimensional vector space.
[0022] When entity and relationship capture fails, re-collect and analyze text data;
[0023] When entities and relationships are captured successfully, a knowledge graph is constructed based on the captured entities and relationships, and a large language model is trained to obtain a trained large language model.
[0024] As a preferred solution of the text processing method based on a large language model described in the present invention, an adaptive task selector is used to automatically match the pre-training data set to the downstream task type, and at the same time, an incremental learning framework and a meta-learning algorithm are used to optimize hyperparameters, including the following steps:
[0025] Define downstream task types, and define feature sets and evaluation metrics for each task type;
[0026] Use a meta-learning algorithm to train an adaptive task selector based on the defined downstream task type, input a pre-trained dataset, and output the predicted downstream task type and hyperparameter configuration recommendations;
[0027] The latest text data is regularly obtained from information sources, added to the parallel corpus after cleaning and annotation, and merged with historical text data to form a new pre-training dataset. The small batch gradient descent optimization method is combined with the prediction results to update some hyperparameters of the trained large language model.
[0028] As a preferred solution of the text processing method based on a large language model described in the present invention, wherein: considering non-text data, using a multimodal model to capture different types of context clues, and using a cross-cultural cognitive framework to understand expressions, the following steps are included:
[0029] Collect non-text data related to the target domain;
[0030] Using a multimodal deep learning architecture, text data and non-text data are projected into the same dimension and jointly represented in a unified vector space to form multimodal data;
[0031] Pair multimodal data with text data to form joint training samples, and use an end-to-end approach to simultaneously train the text encoder and the encoder of other modalities;
[0032] Using conditional random field technology, we add a cross-cultural adaptation layer to the large language model to perform cultural sensitivity analysis on the input text and identify the cultural features involved;
[0033] Use parallel corpora from multiple languages and cultures for pre-training, so that the large language model can initially master the expressions of different cultures and obtain an updated large language model.
[0034] As a preferred solution of the text processing method based on the large language model described in the present invention, wherein: based on the updated large language model, a confidence evaluation method is established, and a high-confidence text processing result is output, including the following steps:
[0035] Define confidence metrics;
[0036] The confidence index is Z-Score standardized to a numerical value, and the confidence score is calculated based on the confidence index, expressed as,
[0037]
[0038] Among them, S represents the confidence score, w i represents the dynamic weight of the i-th confidence indicator, x irepresents the standardized score of the i-th confidence indicator, i represents the index of the confidence indicator, and n represents the number of confidence indicators;
[0039] According to historical text data, set the confidence threshold;
[0040] When the confidence score is less than or equal to the confidence threshold, the prediction result is judged to be low confidence and marked for further review;
[0041] When the confidence score is greater than the confidence threshold, the prediction result is judged to be high confidence;
[0042] Arrange all high-confidence prediction results in descending order and output the text processing result with the highest confidence score.
[0043] As a preferred solution of the text processing method based on a large language model described in the present invention, wherein: according to the high-confidence text processing results, a personalized user portrait is constructed and the recommended content is optimized, including the following steps:
[0044] Collect user interaction data from multiple channels to form user activity logs;
[0045] Use natural language processing technology to extract personalized features from user activity logs;
[0046] Build a multi-dimensional user portrait based on the extracted personalized features;
[0047] Based on the high-confidence text processing results, mark the interest tags in the user portrait;
[0048] Provide users with customized recommended content based on the interest tags in the user portrait.
[0049] In a second aspect, the present invention provides a text processing system based on a large language model, comprising:
[0050] The pre-training module is responsible for collecting text data and generating variant samples through data enhancement technology to obtain a pre-training data set, introduce the knowledge graph, and output the trained large language model;
[0051] The fine-tuning module is responsible for automatically matching the pre-training dataset to the downstream task type using an adaptive task selector. At the same time, it uses an incremental learning framework and meta-learning algorithm to optimize hyperparameters.
[0052] The fusion module is responsible for considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions;
[0053] The evaluation module is responsible for establishing a confidence evaluation method based on the updated large language model and outputting high-confidence text processing results;
[0054] The personalization module is responsible for building personalized user portraits and optimizing recommended content based on high-confidence text processing results.
[0055] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the text processing method based on a large language model as described in the first aspect of the present invention is implemented.
[0056] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the text processing method based on a large language model as described in the first aspect of the present invention.
[0057] The beneficial effects of the present invention are as follows: by introducing a multimodal model and a cross-cultural cognitive framework, the present invention effectively improves the performance of large language models in dealing with complex real-world problems. Specifically, the use of multimodal models can better capture different types of data clues and improve the large language model's ability to understand context; while the cross-cultural cognitive framework enhances the large language model's understanding of expression habits in different cultural and linguistic backgrounds, ensuring its applicability and accuracy in global application scenarios. In addition, by constructing personalized user portraits, the recommended content is optimized, more accurate service provision is achieved, the user experience is significantly improved, and strong support is provided for applications in various fields. This comprehensive improvement not only expands the scope of application of large language models, but also improves their ability to handle fuzzy or ambiguous information, laying a solid foundation for achieving high-quality human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0059] Figure 1 This is a flowchart of the text processing method based on a large language model in Example 1.
[0060] Figure 2 Schematic diagram of a text processing system based on a large language model in Example 1. DETAILED DESCRIPTION
[0061] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0063] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0064] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a text processing method based on a large language model, comprising the following steps:
[0065] S1. Collect text data and generate variant samples through data enhancement technology to obtain pre-trained data sets, introduce knowledge graphs, and output the trained large language model, including the following steps:
[0066] S1.1. Text data include authoritative literature, professional forums and academic journals;
[0067] Authoritative literature: from peer-reviewed professional journals and books, which usually contain the latest research results and technical details;
[0068] Professional Forum: A discussion platform between industry experts and practitioners, providing practical application cases and problem-solving strategies;
[0069] Academic Journals: Research papers published in formal publications that cover a wide range of theoretical foundations and empirical analysis;
[0070] Non-text data includes images and tables.
[0071] S1.2. Determine the target domain and identify the main information sources within the target domain, such as academic databases (such as IEEE Xplore (Chinese: IEEE Exploration Digital Library), PubMed (Chinese: Biomedical Literature Database)), professional websites (such as Stack Overflow (Chinese: Programming Q&A Community), GitHub (Chinese: Code Hosting and Collaboration Platform)), and related social media platforms or forums;
[0072] Use crawler technology to access selected primary information sources and download text data;
[0073] Using the methods of synonym replacement, sentence reorganization and context expansion, variant samples are obtained. Variant samples refer to new data samples generated by making specific transformations or modifications to the original data.
[0074] Synonym replacement: Based on WordNet (lexical network) or other lexical resource libraries, some words in the sentence are randomly replaced with their synonyms, keeping the meaning of the sentence unchanged but adding variants;
[0075] Sentence reorganization: adjust the sentence structure without changing its meaning, such as changing the subject-verb-object order or merging / splitting short sentences, to improve the large language model's ability to understand different forms of expression;
[0076] Context expansion: Add background information or introduce related topics to make sentences closer to real application scenarios, while also training large language models to better capture context associations;
[0077] Integrate text data and variant samples to generate pre-training datasets.
[0078] S1.3. Analyze the unique expression habits, metaphors, and slang of various cultures to form a cultural feature tag library. Use bilingual comparison to build a parallel corpus, that is, different language versions of the same content, to help the large language model learn how to accurately convert and interpret cross-cultural differences;
[0079] It should be noted that bilingual comparison is the presentation of the same document or text paragraph in two different languages, and the contents of the two language versions correspond to each other and have the same meaning.
[0080] The BERT-based NER (named entity recognition based on BERT) model is applied to automatically identify entities in the text, and experts in the target field are invited to review the results of automatic annotation, correct incorrectly identified entities, supplement missed important entities, and annotate entity types (such as people, places, events, organizations, etc.) to provide a basis for subsequent relationship extraction.
[0081] Through dependency syntactic analysis (such as spaCy (spaCy natural language processing library), StanfordNLP (Stanford Natural Language Processing Group Toolkit), etc.), the grammatical structure of the sentence is analyzed to find the potential relationship between entities. For complex relationships, logical reasoning or graph neural networks can be used to further refine and verify them;
[0082] Specifically, the input text is segmented, POS tagged, and named entity recognized to prepare for dependency syntactic analysis, ensuring that each sentence is correctly segmented and all important entities (such as names of people, places, organizations, etc.) are marked. The dependency syntactic analysis tool is run to generate a dependency tree for the sentence, in which nodes represent words and edges represent dependency relationships between words (such as subject-predicate, verb-object, etc.). The dependency tree is analyzed to find dependency paths related to the labeled entities, especially those paths that indicate possible relationships between entities (such as employed by or located in, etc.). Based on the dependency path pattern, possible relationship candidate pairs are extracted. For example, when one entity is found as the subject and the other entity as the object in the dependency tree, and there is a specific type of verb connection between the two, it can be considered that there is a certain relationship.
[0083] The TransE (Translation Embedding) algorithm is used to optimize triple representation and map entities and relations into a low-dimensional vector space.
[0084] It should be noted that the goal of setting the loss function is to minimize the scores of all correct triplets and maximize the scores of incorrect triplets, which can be expressed by the following formula:
[0085]
[0086] Among them, f(h,r,t) represents the score of the triple of the head entity h connected to the tail entity t through the relationship r. The smaller the score, the more likely the triple is to be true. Conversely, the larger the score, the more likely the triple is to be false. h represents the vector representation of the head entity h. Each entity is mapped into a vector space of fixed dimension. This vector captures the semantic features of the entity. R represents the vector representation of the relation r. Similarly, each relation is also mapped into the same vector space to represent the semantic features of the relation. t Represents the vector representation of the tail entity t. Similar to the head entity, the tail entity also has its corresponding vector representation. represents the L2 norm squared, that is, the square of the Euclidean distance.
[0087] In order to speed up the training process and avoid overfitting, negative sampling technology is used, that is, a negative sample is generated by randomly replacing an element from the existing triple. In each iteration, in addition to the positive samples, a certain number of negative samples are also considered to participate in the loss calculation. Stochastic gradient descent or other optimization algorithms are used to update the vector representation of entities and relationships, so that the loss function is gradually reduced.
[0088] When entity and relationship capture fails, re-collect and analyze text data;
[0089] When entities and relationships are captured successfully, a knowledge graph is constructed based on the captured entities and relationships, and the knowledge graph is used to train a large language model. The entities and relationships in the knowledge graph are embedded and loaded into the large language model, so that the large language model can directly access these structured knowledge during the training process, thereby obtaining a trained large language model.
[0090] It should be noted that by accurately locating and acquiring high-quality text data, including authoritative documents, professional forums, and academic journals, and combining data enhancement techniques such as synonym replacement, sentence reorganization, and context expansion, a variety of variant samples were generated. This not only builds a solid data foundation, but also significantly improves the professionalism and accuracy of the large language model. At the same time, by analyzing various cultural characteristics and building parallel corpora, applying the BERT-based NER model for entity recognition, and using the TransE algorithm to optimize triple representation, the large language model's ability to understand language expressions in different cultural backgrounds is enhanced, ensuring its applicability and accuracy in global application scenarios.
[0091] S2, using an adaptive task selector to automatically match the pre-training dataset to the downstream task type, and at the same time, using an incremental learning framework and meta-learning algorithm to optimize hyperparameters, including the following steps:
[0092] S2.1. Define downstream task types, such as text classification, named entity recognition, relation extraction, and question-answering systems, and define feature sets and evaluation metrics for each task type. For example, for text classification tasks, factors such as vocabulary richness and sentence length distribution can be considered;
[0093] Use a meta-learning algorithm to train an adaptive task selector based on the defined downstream task type, input the statistical features of the pre-trained dataset (such as text features, semantic features, and other features), and output the predicted downstream task type and hyperparameter configuration recommendations, such as learning rate, batch size, and regularization coefficient;
[0094] It should be noted that text features include vocabulary size: counting the number of different words in the data set to reflect the diversity of the language; syntactic structure complexity: parsing sentence structure through dependency syntax analysis tools (such as spaCy and StanfordNLP) to calculate indicators such as average sentence length and proportion of complex clauses; vocabulary richness: measuring the diversity of vocabulary, which can be quantified using word frequency distribution or TF-IDF value; topic distribution: using topic models (such as LDA) to analyze the topic distribution of documents to understand the main topic areas covered by the data set. Semantic features include sentiment tendency: using sentiment analysis tools to evaluate the emotional color of the text (such as positive, negative and neutral); entity density: calculating the number and types of entities appearing in the text, reflecting the professionalism and information density of the content. Other features include domain-specific terminology: identifying and counting the frequency of professional terms in the field to help determine the industry background of the task; data set size: recording the total number of samples in the data set, which affects the training time and resource requirements of large language models. Train the adaptive task selector, specifically construct the statistical features of the pre-trained dataset into a feature vector as the input of the large language model. The output of the large language model should include two parts: one is the predicted type that is most suitable for the downstream task; the other is the hyperparameter configuration recommendation for the downstream task type. This can be achieved through multi-label classification or multi-output regression. Define an appropriate loss function, combining the accuracy of the downstream task type prediction (such as cross entropy loss) and the effect of the hyperparameter configuration (such as the performance gap of the large language model) to ensure that the large language model can find the best trade-off between the two.
[0095] S2.2, regularly obtain the latest text data from information sources, and add it to the parallel corpus after cleaning and annotation;
[0096] Each time a batch of new data is received, it is merged with the historical text data to form a new pre-training dataset;
[0097] The optimization method of mini-batch gradient descent is used in combination with the prediction results to update some hyperparameters of the trained large language model, avoiding comprehensive retraining of the entire large language model and thus saving computing resources.
[0098] It should be noted that by defining a variety of downstream task types and corresponding feature sets and evaluation indicators, and using a meta-learning algorithm to predict the most suitable downstream task type and its hyperparameter configuration recommendations, the task selection process is automated and the hyperparameter settings are precisely adjusted. The latest text data is regularly obtained from the information source, merged with the historical text data to form a new pre-training data set, and some hyperparameters are updated to maintain the timeliness and adaptability of the large language model while saving computing resources. This method not only greatly improves the performance of the large language model on different tasks and shortens the development cycle, but also ensures that the large language model can promptly reflect the latest industry trends, reduce the cost of retraining, and improve resource utilization efficiency.
[0099] S3. Consider non-text data, use multimodal models to capture different types of context clues, and use cross-cultural cognitive frameworks to understand expressions, including the following steps:
[0100] S3.1. Collect non-text data related to the target domain, for example, in the medical field, it may be X-rays and electrocardiograms; in news reports, it may be pictures and video clips;
[0101] For image data, perform preprocessing operations such as resizing, color space conversion, and normalization to make it meet the input requirements of the large language model;
[0102] For audio data, use selected tools (e.g., Affectiva (emotion analysis technology) and BeyondVerbal (speech emotion analysis technology)) or APIs to process the audio data, perform speech recognition to generate transcriptions, and extract acoustic features to capture the emotional color or tone of voice;
[0103] Use multimodal deep learning architectures, such as ViT (Vision Transformer) and CLIP (Contrastive Language-Image Pre-training), to project text data and non-text data into the same dimension, represent them together in a unified vector space, and form multimodal data.
[0104] S3.2. Pair multimodal data with text data to form joint training samples, and use an end-to-end approach to simultaneously train the text encoder and the encoder of other modalities, so that the large language model can automatically select the most relevant modal information to assist decision-making during the inference phase;
[0105] It should be noted that when pairing multimodal data with text data, if both the multimodal data and the text data are accompanied by unique identifiers (such as file names, ID numbers), one-to-one pairing can be performed directly through these identifiers; for time series data (such as frames in a video and corresponding subtitles), timestamp information can be used for accurate pairing; further, pairing is performed by calculating the similarity score between the text data and the multimodal data. For example, use pre-trained text encoders and image encoders to extract feature vectors of text and images respectively, and then calculate cosine similarity or other distance metrics. Create a record for each paired multimodal and text data, containing all relevant information (such as paths, labels, feature vectors, etc.). Ensure that each record is a complete training sample that can be correctly loaded and parsed during the training process to obtain a joint training sample.
[0106] S3.3. Using conditional random field technology, we add a cross-cultural adaptation layer to the large language model, perform cultural sensitivity analysis on the input text, and identify the cultural features involved;
[0107] It should be noted that adding a cross-cultural adaptation layer to the large language model is to perform cultural sensitivity analysis on the input text by integrating conditional random field technology, identify and annotate the cultural features involved, and thus enhance the large language model's ability to understand language expressions in different cultural backgrounds;
[0108] Use parallel corpora from multiple languages and cultures for pre-training, so that the large language model can initially master the expressions of different cultures and obtain an updated large language model. When encountering a new culture or language, it can quickly adapt to the new environment by fine-tuning a small number of parameters without having to retrain the entire large language model from scratch.
[0109] It should be noted that by processing non-text data (such as images and audio) in the target domain, preprocessing operations are performed, and a multimodal deep learning architecture is used to represent text and non-text data together in a unified vector space, the effective fusion of text and non-text data is achieved, different types of data clues are captured, the ability of the large language model to understand the context is enhanced, and more comprehensive information support is provided. The attention mechanism is designed to calculate the similarity scores between different modal features, and the softmax function is used to convert them into attention weights, and the weighted summation of different modal features is performed, which effectively captures the relationship and complementary information between different modalities, improves the performance of the large language model in multimodal tasks, and enhances the accuracy and reliability of decision-making. A cross-cultural adaptation layer is added, and conditional random field technology is used to perform cultural sensitivity analysis on the input text, identify and annotate the cultural features involved, and significantly improve the performance of the large language model in cross-cultural communication, expand its scope of application, and ensure its applicability and accuracy in global application scenarios.
[0110] S4. Based on the updated large language model, a confidence assessment method is established and a high-confidence text processing result is output, including the following steps:
[0111] S4.1. Define confidence metrics, including probability scores, consistency check results, and contextual relevance;
[0112] Probability score: For classification tasks, the probability value output by the large language model is directly used as the preliminary confidence indicator;
[0113] Consistency check results: By comparing the prediction results of multiple models (such as different base learners in ensemble learning) or the same model under different hyperparameters, their consistency is evaluated. Consistent results usually mean higher confidence.
[0114] Contextual relevance: Analyze the relevance of the prediction results to the input text and its context to ensure that the output of the large language model conforms to the contextual logic.
[0115] S4.2, the confidence index is Z-Score standardized to a numerical value, and the confidence score is calculated based on the confidence index, expressed as,
[0116]
[0117] Among them, S represents the confidence score, w i represents the dynamic weight of the i-th confidence indicator, which is determined by the entropy value, x i represents the standardized score of the i-th confidence indicator, i represents the index of the confidence indicator, and n represents the number of confidence indicators.
[0118] S4.3. Set a confidence threshold based on historical text data;
[0119] When the confidence score is less than or equal to the confidence threshold, the prediction result is judged to be low confidence and marked for further review;
[0120] When the confidence score is greater than the confidence threshold, the prediction result is judged to be high confidence;
[0121] Arrange all high-confidence prediction results in descending order and output the text processing result with the highest confidence score.
[0122] It should be noted that by defining three confidence indicators, namely probability score, consistency check result and context relevance, and standardizing these indicators, a confidence evaluation system is established to evaluate the reliability of the prediction results of the large language model, ensuring the high quality and reliability of the output results and reducing the false alarm rate. The confidence score is calculated according to the confidence indicator, and a reasonable confidence threshold is set according to historical data, which enables accurate evaluation of the prediction results, distinguishes between high-confidence and low-confidence results, ensures that only high-confidence results are adopted, and reduces the risk of wrong decisions. The high-confidence prediction results are sorted in descending order, and the text processing results with the highest confidence are output, ensuring the optimality and reliability of the final output results, significantly improving user satisfaction, providing more accurate services, and enhancing user trust.
[0123] S5. Based on the high-confidence text processing results, build a personalized user portrait and optimize the recommended content, including the following steps:
[0124] Collect user interaction data from multiple channels, including but not limited to historical query records, click behaviors, comment content, etc., to form user activity logs;
[0125] Use natural language processing technology to extract personalized features from user activity logs, such as query topic distribution, common word frequency, and preferred areas;
[0126] Based on the extracted personalized features, a multi-dimensional user portrait is constructed, covering basic information (age, gender, etc.), interests and hobbies, professional skills and other aspects. Clustering algorithms (such as K-means (K-means clustering algorithm) and DBSCAN (density-based spatial clustering algorithm)) are introduced to segment the user group and identify subgroups with similar features, thereby achieving more accurate personalized services;
[0127] Specifically, building a multi-dimensional user portrait is through analyzing and extracting the user's personalized characteristics, such as interest preferences, behavior patterns and background information, and then using these characteristics to build a comprehensive user portrait covering basic information, interests and hobbies, professional skills and other aspects to achieve accurate personalized services.
[0128] Based on the high-confidence text processing results, mark the interest tags in the user portrait;
[0129] It should be noted that sentiment analysis tools are used to evaluate users' sentiment attitudes towards different topics (such as positive, negative, and neutral). For topics with positive evaluations, the corresponding interest tag weights are increased; otherwise, the weights are reduced; users' questions or comments are analyzed to infer their true intentions (such as seeking advice, expressing opinions, sharing experiences, etc.), and the allocation of interest tags is adjusted accordingly. Based on the defined tag mapping rules and the extracted personalized features, interest tags are automatically added to each user portrait.
[0130] Provide users with customized recommended content, such as articles, products, services, etc., based on the interest tags in the user portrait.
[0131] It should be noted that by collecting user interaction data and extracting personalized features, such as query topic distribution, common word frequency, and preferred fields, a user portrait covering basic information, hobbies, and professional skills is constructed, and recommended content is optimized based on the interest tags in the user portrait. This method forms a detailed user portrait, provides customized recommendation services based on user preferences, provides users with a more accurate and personalized service experience, meets users' unique needs, significantly improves user participation and satisfaction, and enhances the platform's attractiveness and user stickiness. By analyzing and extracting users' personalized features, accurate personalized services are achieved, further optimizing the user experience, and improving user loyalty and the platform's market competitiveness.
[0132] This embodiment also provides a text processing system based on a large language model, including:
[0133] The pre-training module is responsible for collecting text data and generating variant samples through data enhancement technology to obtain a pre-training data set, introduce the knowledge graph, and output the trained large language model;
[0134] The fine-tuning module is responsible for automatically matching the pre-training dataset to the downstream task type using an adaptive task selector. At the same time, it uses an incremental learning framework and meta-learning algorithm to optimize hyperparameters.
[0135] The fusion module is responsible for considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions;
[0136] The evaluation module is responsible for establishing a confidence evaluation method based on the updated large language model and outputting high-confidence text processing results;
[0137] The personalization module is responsible for building personalized user portraits and optimizing recommended content based on high-confidence text processing results.
[0138] This embodiment also provides a computer device, which is suitable for the case of a text processing method based on a large language model, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the text processing method based on a large language model proposed in the above embodiment.
[0139] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0140] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the text processing method based on a large language model as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0141] In summary, the beneficial effects of the present invention are as follows: by introducing a multimodal model and a cross-cultural cognitive framework, the present invention effectively improves the performance of large language models in dealing with complex real-world problems. Specifically, the use of multimodal models can better capture different types of data clues and improve the large language model's ability to understand context; while the cross-cultural cognitive framework enhances the large language model's understanding of expression habits in different cultural and linguistic backgrounds, ensuring its applicability and accuracy in global application scenarios. In addition, by constructing personalized user portraits, the recommended content is optimized, more accurate service provision is achieved, the user experience is significantly improved, and strong support is provided for applications in various fields. This comprehensive improvement not only expands the scope of application of large language models, but also improves their ability to handle fuzzy or ambiguous information, laying a solid foundation for achieving high-quality human-computer interaction.
[0142] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A text processing method based on a large language model, characterized in that: include, Collect text data and generate variant samples through data augmentation technology to obtain pre-training data sets, introduce knowledge graphs, and output the trained large language model; Adopting adaptive task selector, the pre-training dataset is automatically matched to the downstream task type. At the same time, the incremental learning framework and meta-learning algorithm are used to optimize hyperparameters. Consider non-textual data, use multimodal models to capture different types of contextual clues, and use cross-cultural cognitive frameworks to understand expressions; Based on the updated large language model, a confidence assessment method is established and high-confidence text processing results are output; Based on high-confidence text processing results, build personalized user portraits and optimize recommended content.
2. The text processing method based on a large language model as claimed in claim 1, characterized in that: The text data includes authoritative documents, professional forums and academic journals; The non-text data includes images and tables.
3. The text processing method based on a large language model as claimed in claim 2, characterized in that: Collect text data, generate variant samples through data enhancement technology, obtain pre-training data set, introduce knowledge graph, and output the trained large language model, including the following steps: Determine the target area and identify the main information sources within the target area; Use crawler technology to access selected primary information sources and download text data; Use synonym replacement, sentence reorganization and context expansion to obtain variant samples; Integrate text data and variant samples to generate pre-training datasets; Analyze the unique expression habits, metaphors and slang of various cultures to form a cultural feature label library, and use bilingual comparison to build a parallel corpus; Apply the BERT-based NER model to automatically identify entities in text and label entity types; Through dependency parsing, the grammatical structure of the sentence is analyzed to find the potential relationship between entities; The TransE algorithm is used to optimize triple representation and map entities and relations into low-dimensional vector space. When entity and relationship capture fails, re-collect and analyze text data; When entities and relationships are captured successfully, a knowledge graph is constructed based on the captured entities and relationships, and a large language model is trained to obtain a trained large language model.
4. The text processing method based on a large language model as claimed in claim 3, characterized in that: Adopting adaptive task selector, the pre-training data set is automatically matched to the downstream task type. At the same time, the incremental learning framework and meta-learning algorithm are used to optimize the hyperparameters, including the following steps: Define downstream task types, and define feature sets and evaluation metrics for each task type; Use a meta-learning algorithm to train an adaptive task selector based on the defined downstream task type, input a pre-trained dataset, and output the predicted downstream task type and hyperparameter configuration recommendations; The latest text data is regularly obtained from information sources, added to the parallel corpus after cleaning and annotation, and merged with historical text data to form a new pre-training dataset. The small batch gradient descent optimization method is combined with the prediction results to update some hyperparameters of the trained large language model.
5. The text processing method based on a large language model as claimed in claim 4, characterized in that: Considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions, including the following steps, Collect non-text data related to the target domain; Using a multimodal deep learning architecture, text data and non-text data are projected into the same dimension and jointly represented in a unified vector space to form multimodal data; Pair multimodal data with text data to form joint training samples, and use an end-to-end approach to simultaneously train the text encoder and the encoder of other modalities; Using conditional random field technology, we add a cross-cultural adaptation layer to the large language model to perform cultural sensitivity analysis on the input text and identify the cultural features involved; Use parallel corpora from multiple languages and cultures for pre-training, so that the large language model can initially master the expressions of different cultures and obtain an updated large language model.
6. The text processing method based on a large language model as claimed in claim 5, characterized in that: Based on the updated large language model, a confidence assessment method is established and high-confidence text processing results are output, including the following steps: Define confidence metrics; The confidence index is Z-Score standardized to a numerical value, and the confidence score is calculated based on the confidence index, expressed as, Among them, S represents the confidence score, w i represents the dynamic weight of the i-th confidence indicator, x i represents the standardized score of the i-th confidence indicator, i represents the index of the confidence indicator, and n represents the number of confidence indicators; According to historical text data, set the confidence threshold; When the confidence score is less than or equal to the confidence threshold, the prediction result is judged to be low confidence and marked for further review; When the confidence score is greater than the confidence threshold, the prediction result is judged to be high confidence; Arrange all high-confidence prediction results in descending order and output the text processing result with the highest confidence score.
7. The text processing method based on a large language model as claimed in claim 6, characterized in that: Based on the high-confidence text processing results, build a personalized user portrait and optimize the recommended content, including the following steps: Collect user interaction data from multiple channels to form user activity logs; Use natural language processing technology to extract personalized features from user activity logs; Build a multi-dimensional user portrait based on the extracted personalized features; Based on high-confidence text processing results, mark interest tags in user portraits; Provide users with customized recommended content based on the interest tags in the user portrait.
8. A text processing system based on a large language model, based on the text processing method based on a large language model according to any one of claims 1 to 7, characterized in that: include, The pre-training module is responsible for collecting text data and generating variant samples through data enhancement technology to obtain a pre-training data set, introduce the knowledge graph, and output the trained large language model; The fine-tuning module is responsible for automatically matching the pre-training dataset to the downstream task type using an adaptive task selector. At the same time, it uses an incremental learning framework and meta-learning algorithm to optimize hyperparameters. The fusion module is responsible for considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions; The evaluation module is responsible for establishing a confidence evaluation method based on the updated large language model and outputting high-confidence text processing results; The personalization module is responsible for building personalized user portraits and optimizing recommended content based on high-confidence text processing results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the text processing method based on a large language model described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the text processing method based on a large language model described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Text translation model training method and device and text translation method and device
CN117094418A
Question and answer processing method and device, storage medium and program product
CN118939780A
Dynamic Semiotic Systemic Knowledge Compiler System and Methods
US20160342398A1
Cited By
Multi-modal task text processing optimization method based on large language model
CN121234902A
Semantic analysis and recognition method based on artificial intelligence
CN121257548A