A text processing method and system based on large language model

By combining data enhancement and multimodal models with a cross-cultural cognitive framework, we solved the adaptability problem of large language models in multimodal and cross-cultural scenarios, achieved efficient text processing and personalized services, and improved the applicability and accuracy of the model.

CN120104781BActive Publication Date: 2025-09-12BEIJING TAIHE GUANFU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510099874.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-09-12
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing large language models have poor adaptability in processing multimodal information and cross-cultural scenarios. They find it difficult to capture contextual clues of non-text data such as images and tables, and their adaptive task selection mechanism cannot dynamically respond to changes in dataset characteristics, limiting their flexibility and efficiency.

Method used

Variant samples are generated through data augmentation technology, knowledge graphs are introduced, multimodal models and cross-cultural cognitive frameworks are utilized, and adaptive task selectors, incremental learning frameworks and meta-learning algorithms are combined to optimize hyperparameters, establish confidence assessment methods, build personalized user portraits, and optimize recommended content.

Benefits of technology

It improves the performance of large language models in dealing with complex real-world problems, enhances the ability to understand context and cross-cultural expression habits, ensures applicability and accuracy in global application scenarios, and realizes precise personalized services and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104781B_ABST
    Figure CN120104781B_ABST
Patent Text Reader

Abstract

The present invention discloses a text processing method and system based on a large language model, which relates to the field of computer technology. The method and system include collecting text data, generating variant samples through data enhancement technology, obtaining a pre-training data set, introducing a knowledge graph, and outputting a trained large language model; utilizing an adaptive task selector, an incremental learning framework, and a meta-learning algorithm to optimize hyperparameters; considering non-text data, utilizing a multimodal model to capture different types of context clues, and utilizing a cross-cultural cognitive framework to understand expressions; establishing a confidence assessment method based on the updated large language model, and outputting a high-confidence text processing result; constructing a personalized user portrait based on the high-confidence text processing result, and optimizing recommended content. The present invention effectively improves the performance of the large language model in processing complex real-world problems by introducing a multimodal model and a cross-cultural cognitive framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a text processing method and system based on a large language model. Background Art

[0002] In recent years, with the rapid development of natural language processing (NLP) technology, text processing methods based on large language models have gradually become a research hotspot. In existing technologies, pre-trained language models such as BERT and RoBERTa have achieved remarkable results on general corpora and are widely used in various downstream tasks. However, the application of these models in specific fields or cross-cultural scenarios still faces challenges. On the one hand, traditional methods have limited processing of non-text data and find it difficult to capture the rich contextual clues brought by multimodal information such as images and tables; on the other hand, the lack of understanding of the expression habits and metaphors of different cultures leads to poor adaptability of the model in global application scenarios. In addition, existing adaptive task selection mechanisms usually rely on static configurations and cannot dynamically respond to changes in dataset characteristics, which limits their flexibility and efficiency. Summary of the Invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] Therefore, the present invention provides a text processing method based on a large language model to solve the problems of multimodal data processing and insufficient cross-cultural understanding.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] In a first aspect, the present invention provides a text processing method based on a large language model, which includes collecting text data, generating variant samples through data augmentation technology, obtaining a pre-training dataset, introducing a knowledge graph, and outputting a trained large language model;

[0007] Adopting an adaptive task selector, the pre-training dataset is automatically matched to the downstream task type. At the same time, an incremental learning framework and meta-learning algorithm are used to optimize hyperparameters.

[0008] Consider non-textual data, use multimodal models to capture different types of contextual clues, and utilize cross-cultural cognitive frameworks to understand expressions;

[0009] Based on the updated large language model, a confidence assessment method is established and high-confidence text processing results are output;

[0010] Based on high-confidence text processing results, build personalized user portraits and optimize recommended content.

[0011] As a preferred solution of the text processing method based on a large language model of the present invention, wherein: the text data includes authoritative documents, professional forums and academic journals;

[0012] The non-text data includes images and tables.

[0013] As a preferred solution of the text processing method based on the large language model described in the present invention, the method includes the following steps: collecting text data, generating variant samples through data enhancement technology, obtaining a pre-training data set, introducing a knowledge graph, and outputting a trained large language model.

[0014] Determine the target area and identify the main information sources within the target area;

[0015] Use crawler technology to access selected primary information sources and download text data;

[0016] Use synonym replacement, sentence reorganization and context expansion methods to obtain variant samples;

[0017] Integrate text data and variant samples to generate pre-training datasets;

[0018] Analyze the unique expressions, metaphors, and slang of various cultures to form a cultural feature tag library, and use bilingual comparison to build a parallel corpus;

[0019] Apply the BERT-based NER model to automatically identify entities in text and label entity types;

[0020] Through dependency parsing, the grammatical structure of the sentence is analyzed to find the potential relationship between entities;

[0021] Use the TransE algorithm to optimize triple representation and map entities and relations into a low-dimensional vector space;

[0022] When entity and relationship capture fails, re-collect and analyze text data;

[0023] When entities and relationships are captured successfully, a knowledge graph is constructed based on the captured entities and relationships, and a large language model is trained to obtain a trained large language model.

[0024] As a preferred solution of the text processing method based on a large language model described in the present invention, an adaptive task selector is used to automatically match the pre-training dataset to the downstream task type. At the same time, an incremental learning framework and a meta-learning algorithm are used to optimize hyperparameters, including the following steps:

[0025] Define downstream task types and define feature sets and evaluation metrics for each task type;

[0026] Use a meta-learning algorithm to train an adaptive task selector based on the defined downstream task type, input a pre-training dataset, and output the predicted downstream task type and hyperparameter configuration recommendations;

[0027] The latest text data is regularly obtained from information sources, and after cleaning and annotation, it is added to the parallel corpus, which is then merged with historical text data to form a new pre-training dataset. The small-batch gradient descent optimization method is combined with the prediction results to update some hyperparameters of the trained large language model.

[0028] As a preferred solution of the text processing method based on the large language model described in the present invention, wherein: considering non-text data, using a multimodal model to capture different types of context clues, and using a cross-cultural cognitive framework to understand expressions, the following steps are included:

[0029] Collect non-textual data related to the target domain;

[0030] Using a multimodal deep learning architecture, text data and non-text data are projected into the same dimension and jointly represented in a unified vector space to form multimodal data;

[0031] Pair multimodal data with text data to form joint training samples, and use an end-to-end approach to simultaneously train the text encoder and the encoder of other modalities;

[0032] Using conditional random field technology, we add a cross-cultural adaptation layer to the large language model to perform cultural sensitivity analysis on the input text and identify the cultural characteristics involved;

[0033] Use parallel corpora from multiple languages ​​and cultures for pre-training, so that the large language model can initially grasp the expressions of different cultures and obtain an updated large language model.

[0034] As a preferred solution of the text processing method based on the large language model described in the present invention, wherein: based on the updated large language model, a confidence assessment method is established and a high-confidence text processing result is output, including the following steps:

[0035] Define confidence metrics;

[0036] The confidence index is Z-Score standardized to a numerical value, and the confidence score is calculated based on the confidence index, which is expressed as,

[0037]

[0038] Among them, S represents the confidence score, w i represents the dynamic weight of the i-th confidence indicator, x irepresents the standardized score of the i-th confidence indicator, i represents the index of the confidence indicator, and n represents the number of confidence indicators;

[0039] Set confidence thresholds based on historical text data;

[0040] When the confidence score is less than or equal to the confidence threshold, the prediction result is judged to be low confidence and marked for further review;

[0041] When the confidence score is greater than the confidence threshold, the prediction result is judged to be high confidence;

[0042] Arrange all high-confidence prediction results in descending order and output the text processing result with the highest confidence score.

[0043] As a preferred solution of the text processing method based on the large language model described in the present invention, wherein: according to the high-confidence text processing results, a personalized user portrait is constructed and the recommended content is optimized, including the following steps:

[0044] Collect user interaction data from multiple channels to form user activity logs;

[0045] Use natural language processing technology to extract personalized features from user activity logs;

[0046] Build multi-dimensional user portraits based on the extracted personalized features;

[0047] Based on high-confidence text processing results, mark interest tags in user portraits;

[0048] Provide users with customized recommended content based on the interest tags in their user portraits.

[0049] In a second aspect, the present invention provides a text processing system based on a large language model, comprising:

[0050] The pre-training module is responsible for collecting text data and generating variant samples through data augmentation technology to obtain a pre-training dataset, introduce the knowledge graph, and output the trained large language model;

[0051] The fine-tuning module is responsible for automatically matching the pre-training dataset to the downstream task type using an adaptive task selector. At the same time, it uses an incremental learning framework and meta-learning algorithm to optimize hyperparameters.

[0052] The fusion module is responsible for considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions;

[0053] The evaluation module is responsible for establishing a confidence assessment method based on the updated large language model and outputting high-confidence text processing results;

[0054] The personalization module is responsible for building personalized user portraits and optimizing recommended content based on high-confidence text processing results.

[0055] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the text processing method based on a large language model as described in the first aspect of the present invention is implemented.

[0056] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the text processing method based on a large language model as described in the first aspect of the present invention.

[0057] The beneficial effects of the present invention are as follows: by introducing a multimodal model and a cross-cultural cognitive framework, the present invention effectively improves the performance of large language models in dealing with complex real-world problems. Specifically, the use of multimodal models can better capture different types of data clues, thereby improving the large language model's ability to understand the context; while the cross-cultural cognitive framework enhances the large language model's understanding of expression habits in different cultural and language backgrounds, ensuring its applicability and accuracy in global application scenarios. In addition, by constructing personalized user portraits, the recommended content is optimized, more accurate service provision is achieved, the user experience is significantly improved, and strong support is provided for applications in various fields. This comprehensive improvement not only expands the scope of application of large language models, but also improves their ability to handle fuzzy or ambiguous information, laying a solid foundation for achieving high-quality human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 This is a flowchart of the text processing method based on the large language model in Example 1.

[0060] Figure 2 Schematic diagram of the text processing system based on the large language model in Example 1. DETAILED DESCRIPTION

[0061] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0063] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0064] Example 1, with reference to Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a text processing method based on a large language model, comprising the following steps:

[0065] S1. Collect text data and generate variant samples through data enhancement technology to obtain pre-training data sets, introduce knowledge graphs, and output the trained large language model, including the following steps:

[0066] S1.1. Text data includes authoritative literature, professional forums and academic journals;

[0067] Authoritative literature: from peer-reviewed professional journals and books, which usually contain the latest research results and technical details;

[0068] Professional Forum: A discussion platform between industry experts and practitioners, providing practical application cases and problem-solving strategies;

[0069] Academic Journals: Research papers published in formal publications that cover a wide range of theoretical foundations and empirical analysis;

[0070] Non-text data includes images and tables.

[0071] S1.2. Determine the target domain and identify key information sources within it, such as academic databases (e.g., IEEE Xplore, PubMed), professional websites (e.g., Stack Overflow, GitHub), and relevant social media platforms or forums.

[0072] Use crawler technology to access selected primary information sources and download text data;

[0073] Using synonym replacement, sentence reorganization and context expansion methods, variant samples are obtained. Variant samples refer to new data samples generated by performing specific transformations or modifications on the original data.

[0074] Synonym replacement: Based on WordNet (lexical network) or other lexical resource libraries, some words in the sentence are randomly replaced with their synonyms, keeping the meaning unchanged but adding variations;

[0075] Sentence Restructuring: Adjusting sentence structure without changing its meaning, such as changing the subject-verb-object order or merging / splitting short sentences, improves the large language model's ability to understand different forms of expression;

[0076] Context expansion: Adding background information or introducing related topics makes sentences more relevant to real-world scenarios, while also training large language models to better capture contextual associations.

[0077] Integrate text data and variant samples to generate pre-training datasets.

[0078] S1.3. Analyze the unique expressions, metaphors, and slang of various cultures to form a library of cultural feature tags. Use bilingual comparisons to build parallel corpora, i.e., different language versions of the same content, to help large language models learn how to accurately translate and interpret cross-cultural differences.

[0079] It should be noted that bilingualism is the presentation of the same document or text paragraph in two different languages, and the contents of the two language versions correspond to each other and have the same meaning.

[0080] The BERT-based NER (named entity recognition based on BERT) model is applied to automatically identify entities in the text. Experts in the target field are invited to review the results of the automatic annotation, correct incorrectly identified entities, supplement missed important entities, and annotate entity types (such as people, places, events, organizations, etc.) to provide a basis for subsequent relationship extraction.

[0081] Through dependency parsing (such as spaCy (spaCy Natural Language Processing Library), StanfordNLP (Stanford Natural Language Processing Group Toolkit), etc.), the grammatical structure of sentences is analyzed to find the potential relationships between entities. For complex relationships, logical reasoning or graph neural networks can be used to further refine and verify them.

[0082] Specifically, the input text is segmented, tagged with parts of speech, and recognized as named entities to prepare for dependency parsing. This ensures that each sentence is correctly segmented and all important entities (such as names of people, places, organizations, etc.) are marked. The dependency parser runs a dependency tree for the sentence, where nodes represent words and edges represent dependency relationships between words (such as subject-predicate, verb-object, etc.). The dependency tree is analyzed to find dependency paths related to the labeled entities, especially those that indicate possible relationships between entities (such as employed by or located in, etc.). Based on the dependency path pattern, possible relationship candidate pairs are extracted. For example, when one entity is found as the subject and another entity as the object in the dependency tree, and there is a specific type of verb connection between the two, it can be considered that there is a certain relationship.

[0083] Utilize the TransE (Translation Embedding) algorithm to optimize triple representation and map entities and relations into a low-dimensional vector space;

[0084] It should be noted that the goal of setting the loss function is to minimize the scores of all correct triplets and maximize the scores of incorrect triplets, which can be expressed as follows:

[0085]

[0086] Among them, f(h,r,t) represents the score of the triple of the head entity h connected to the tail entity t through the relationship r. The smaller the score, the more likely the triple is to be true. Conversely, the larger the score, the more likely the triple is to be false. h represents the vector representation of the head entity h. Each entity is mapped into a fixed-dimensional vector space. This vector captures the semantic features of the entity. R represents the vector representation of the relationship r. Similarly, each relationship is also mapped into the same vector space to represent the semantic features of the relationship. t Represents the vector representation of the tail entity t. Similar to the head entity, the tail entity also has its corresponding vector representation. represents the square of the L2 norm, that is, the square of the Euclidean distance.

[0087] To accelerate the training process and avoid overfitting, negative sampling is used. This technique randomly replaces an element in the existing triples to generate negative samples. In each iteration, in addition to the positive samples, a certain number of negative samples are also considered in the loss calculation. Stochastic gradient descent or other optimization algorithms are used to update the vector representations of entities and relationships, so that the loss function is gradually reduced.

[0088] When entity and relationship capture fails, re-collect and analyze text data;

[0089] When entities and relationships are captured successfully, a knowledge graph is constructed based on the captured entities and relationships, and the knowledge graph is used to train a large language model. The entities and relationships in the knowledge graph are embedded and loaded into the large language model, so that the large language model can directly access these structured knowledge during the training process, and obtain the trained large language model.

[0090] It should be noted that by precisely locating and acquiring high-quality text data, including authoritative literature, professional forums, and academic journals, and combining data augmentation techniques such as synonym replacement, sentence reorganization, and contextual expansion, a diverse set of variant samples was generated. This not only established a solid data foundation but also significantly improved the professionalism and accuracy of the large language model. Furthermore, by analyzing various cultural characteristics and constructing parallel corpora, applying the BERT-based NER model for entity recognition, and optimizing triple representations using the TransE algorithm, the large language model's ability to understand language expressions in different cultural contexts was enhanced, ensuring its applicability and accuracy in global application scenarios.

[0091] S2, using an adaptive task selector to automatically match the pre-training dataset to the downstream task type, and at the same time, using an incremental learning framework and meta-learning algorithm to perform hyperparameter optimization, including the following steps:

[0092] S2.1. Define downstream task types, such as text classification, named entity recognition, relation extraction, and question-answering systems, and define feature sets and evaluation metrics for each task type. For example, for text classification tasks, factors such as vocabulary richness and sentence length distribution can be considered.

[0093] Using a meta-learning algorithm, an adaptive task selector is trained based on the defined downstream task type. The selected task selector takes in the statistical features of the pre-training dataset (such as text features, semantic features, and other features) and outputs the predicted downstream task type and hyperparameter configuration recommendations, such as learning rate, batch size, and regularization coefficient.

[0094] It should be noted that text features include vocabulary size: counting the number of different words in the data set to reflect the diversity of the language; syntactic structure complexity: parsing sentence structure through dependency syntax analysis tools (such as spaCy and StanfordNLP) to calculate indicators such as average sentence length and the proportion of complex clauses; vocabulary richness: measuring the diversity of vocabulary, which can be quantified using word frequency distribution or TF-IDF value; topic distribution: using topic models (such as LDA) to analyze the topic distribution of documents to understand the main topic areas covered by the data set. Semantic features include sentiment tendency: using sentiment analysis tools to evaluate the emotional color of the text (such as positive, negative and neutral); entity density: calculating the number and types of entities appearing in the text to reflect the professionalism and information density of the content. Other features include domain-specific terminology: identifying and counting the frequency of professional terms in the field to help determine the industry background of the task; dataset size: recording the total number of samples in the dataset, which affects the training time and resource requirements of large language models. Training an adaptive task selector involves constructing a feature vector from the statistical features of the pre-trained dataset as input to the large language model. The output of the large language model should include two parts: first, the predicted type that is most suitable for the downstream task; second, the hyperparameter configuration recommendations for that downstream task type. This can be achieved through multi-label classification or multi-output regression. An appropriate loss function is defined that combines the accuracy of downstream task type predictions (such as cross-entropy loss) and the effect of hyperparameter configuration (such as the performance gap of the large language model) to ensure that the large language model can find the optimal trade-off between the two.

[0095] S2.2. Regularly obtain the latest text data from information sources, clean and annotate it, and add it to the parallel corpus;

[0096] Each time a new batch of data is received, it is merged with the historical text data to form a new pre-training dataset;

[0097] The optimization method of mini-batch gradient descent is combined with the prediction results to update some hyperparameters of the trained large language model, avoiding full retraining of the entire large language model and thus saving computing resources.

[0098] It should be noted that by defining multiple downstream task types and corresponding feature sets and evaluation indicators, and using a meta-learning algorithm to predict the most suitable downstream task type and its hyperparameter configuration recommendations, the task selection process is automated and the hyperparameter settings are precisely adjusted. The latest text data is regularly obtained from information sources, merged with historical text data to form a new pre-training dataset, and some hyperparameters are updated, maintaining the timeliness and adaptability of the large language model while saving computing resources. This approach not only significantly improves the performance of the large language model on different tasks and shortens the development cycle, but also ensures that the large language model can promptly reflect the latest industry trends, reduces the cost of retraining, and improves resource utilization efficiency.

[0099] S3. Consider non-text data, use multimodal models to capture different types of context clues, and use cross-cultural cognitive frameworks to understand expressions, including the following steps:

[0100] S3.1. Collect non-text data related to the target domain. For example, in the medical field, this may be X-rays and electrocardiograms; in news reports, this may be pictures and video clips.

[0101] For image data, perform preprocessing operations such as resizing, color space conversion, and normalization to make it meet the input requirements of large language models;

[0102] For audio data, use selected tools (for example, Affectiva (emotion analysis technology) and BeyondVerbal (speech emotion analysis technology)) or APIs to process the audio data, perform speech recognition to generate transcripts, and extract acoustic features to capture the emotional color or intonation of the voice;

[0103] Using multimodal deep learning architectures such as ViT (Vision Transformer) and CLIP (Contrastive Language-Image Pre-training), text data and non-text data are projected into the same dimension and jointly represented in a unified vector space to form multimodal data.

[0104] S3.2. Pair multimodal data with text data to form joint training samples. Use an end-to-end approach to train both the text encoder and the encoder for the other modality. This allows the large language model to automatically select the most relevant modal information to aid decision-making during inference.

[0105] It should be noted that when pairing multimodal data with text data, if both the multimodal data and the text data are accompanied by unique identifiers (such as file names, ID numbers), one-to-one pairing can be performed directly through these identifiers; for time series data (such as frames in a video and corresponding subtitles), timestamp information can be used for accurate pairing; further, pairing can be performed by calculating the similarity score between the text data and the multimodal data. For example, use pre-trained text encoders and image encoders to extract feature vectors of text and images respectively, and then calculate cosine similarity or other distance metrics. Create a record for each paired multimodal and text data, containing all relevant information (such as paths, labels, feature vectors, etc.). Ensure that each record is a complete training sample that can be correctly loaded and parsed during the training process to obtain a joint training sample.

[0106] S3.3. Using conditional random fields, we add a cross-cultural adaptation layer to the large language model to perform cultural sensitivity analysis on the input text and identify relevant cultural features.

[0107] It should be noted that adding a cross-cultural adaptation layer to the large language model involves integrating conditional random field technology to perform cultural sensitivity analysis on the input text, identify and annotate the relevant cultural features, and thus enhance the large language model's ability to understand language expressions in different cultural backgrounds.

[0108] Pre-training with parallel corpora from multiple languages ​​and cultures allows the large language model to initially master the expressions of different cultures, resulting in an updated large language model. When encountering a new culture or language, it can quickly adapt to the new environment by fine-tuning a small number of parameters without having to retrain the entire large language model from scratch.

[0109] It should be noted that by processing non-text data in the target domain (such as images and audio), performing preprocessing operations, and using a multimodal deep learning architecture to represent text and non-text data together in a unified vector space, the effective fusion of text and non-text data is achieved, capturing different types of data clues, enhancing the large language model's ability to understand context, and providing more comprehensive information support. An attention mechanism is designed to calculate the similarity scores between different modal features and convert them into attention weights using the softmax function. A weighted summation of the different modal features is performed, effectively capturing the relationships and complementary information between the different modalities, improving the performance of the large language model in multimodal tasks, and enhancing the accuracy and reliability of decision-making. A cross-cultural adaptation layer is added, using conditional random field technology to perform cultural sensitivity analysis on the input text, identifying and annotating the cultural features involved. This significantly improves the performance of the large language model in cross-cultural communication, expands its application scope, and ensures its applicability and accuracy in global application scenarios.

[0110] S4. Based on the updated large language model, establish a confidence assessment method and output high-confidence text processing results, including the following steps:

[0111] S4.1. Define confidence metrics, including probability scores, consistency check results, and contextual relevance.

[0112] Probability score: For classification tasks, the probability value output by the large language model is directly used as the preliminary confidence indicator;

[0113] Consistency check results: By comparing the prediction results of multiple models (such as different base learners in ensemble learning) or the same model with different hyperparameters, their consistency is evaluated. Consistent results usually mean higher confidence.

[0114] Contextual relevance: Analyze the relevance of prediction results to the input text and its context to ensure that the output of the large language model conforms to the contextual logic.

[0115] S4.2. Normalize the confidence index to a numerical value using Z-Score, and calculate the confidence score based on the confidence index, expressed as,

[0116]

[0117] Among them, S represents the confidence score, w i represents the dynamic weight of the i-th confidence indicator, which is determined by the entropy value, x i represents the normalized score of the i-th confidence indicator, i represents the index of the confidence indicator, and n represents the number of confidence indicators.

[0118] S4.3. Set a confidence threshold based on historical text data;

[0119] When the confidence score is less than or equal to the confidence threshold, the prediction result is judged to be low confidence and marked for further review;

[0120] When the confidence score is greater than the confidence threshold, the prediction result is judged to be high confidence;

[0121] Arrange all high-confidence prediction results in descending order and output the text processing result with the highest confidence score.

[0122] It should be noted that by defining three confidence indicators, namely probability score, consistency check result, and context relevance, and standardizing these indicators, a confidence assessment system was established to evaluate the reliability of the prediction results of large language models, ensuring the high quality and reliability of the output results and reducing the false alarm rate. By calculating the confidence score based on the confidence indicator and setting a reasonable confidence threshold based on historical data, an accurate assessment of the prediction results was achieved, distinguishing between high-confidence and low-confidence results, ensuring that only high-confidence results are adopted, and reducing the risk of incorrect decisions. High-confidence prediction results are sorted in descending order, and the text processing result with the highest confidence is output, ensuring the optimality and reliability of the final output results, significantly improving user satisfaction, providing more accurate services, and enhancing user trust.

[0123] S5. Based on the high-confidence text processing results, build a personalized user portrait and optimize the recommended content, including the following steps:

[0124] Collect user interaction data from multiple channels, including but not limited to historical query records, click behaviors, comment content, etc., to form user activity logs;

[0125] Use natural language processing technology to extract personalized features from user activity logs, such as query topic distribution, common word frequency, and preferred areas;

[0126] Based on the extracted personalized features, a multi-dimensional user profile is constructed, covering basic information (age, gender, etc.), interests, hobbies, and professional skills. Clustering algorithms (such as K-means and DBSCAN) are introduced to segment user groups and identify subgroups with similar characteristics, thereby achieving more precise personalized services.

[0127] Specifically, building a multi-dimensional user portrait is through analyzing and extracting the user's personalized characteristics, such as interest preferences, behavioral patterns and background information, and then using these characteristics to build a comprehensive user portrait covering basic information, interests and hobbies, professional skills and other aspects to achieve accurate personalized services.

[0128] Based on high-confidence text processing results, mark interest tags in user portraits;

[0129] It's important to note that sentiment analysis tools assess users' emotional attitudes toward different topics (e.g., positive, negative, neutral). For topics with positive evaluations, the corresponding interest tag weight is increased; for topics with negative evaluations, the weight is decreased. Users' questions or comments are analyzed to infer their true intentions (e.g., seeking advice, expressing opinions, sharing experiences, etc.), and interest tag assignments are adjusted accordingly. Based on defined tag mapping rules and extracted personalized features, interest tags are automatically assigned to each user profile.

[0130] Provide users with customized recommended content, such as articles, products, services, etc., based on the interest tags in the user portrait.

[0131] It should be noted that by collecting user interaction data and extracting personalized features such as query topic distribution, common word frequency, and preferred areas, we construct user profiles covering basic information, interests, hobbies, and professional skills. Recommendations are then optimized based on the interest tags in the user profiles. This approach forms detailed user profiles and provides customized recommendation services based on user preferences, providing users with a more precise and personalized service experience that meets their unique needs. This significantly increases user engagement and satisfaction, and enhances the platform's appeal and user stickiness. By analyzing and extracting users' personalized features, we achieve precise personalized services, further optimize the user experience, and improve user loyalty and the platform's market competitiveness.

[0132] This embodiment further provides a text processing system based on a large language model, including:

[0133] The pre-training module is responsible for collecting text data and generating variant samples through data augmentation technology to obtain a pre-training dataset, introduce the knowledge graph, and output the trained large language model;

[0134] The fine-tuning module is responsible for automatically matching the pre-training dataset to the downstream task type using an adaptive task selector. At the same time, it uses an incremental learning framework and meta-learning algorithm to optimize hyperparameters.

[0135] The fusion module is responsible for considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions;

[0136] The evaluation module is responsible for establishing a confidence assessment method based on the updated large language model and outputting high-confidence text processing results;

[0137] The personalization module is responsible for building personalized user portraits and optimizing recommended content based on high-confidence text processing results.

[0138] This embodiment also provides a computer device suitable for the case of a text processing method based on a large language model, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the text processing method based on a large language model proposed in the above embodiment.

[0139] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0140] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the text processing method based on a large language model as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0141] In summary, the beneficial effects of the present invention are as follows: by introducing a multimodal model and a cross-cultural cognitive framework, the present invention effectively improves the performance of large language models in dealing with complex real-world problems. Specifically, the use of multimodal models can better capture different types of data clues, thereby improving the large language model's ability to understand the context; while the cross-cultural cognitive framework enhances the large language model's understanding of expression habits in different cultural and language backgrounds, ensuring its applicability and accuracy in global application scenarios. In addition, by constructing personalized user portraits, the recommended content is optimized, more accurate service provision is achieved, the user experience is significantly improved, and strong support is provided for applications in various fields. This comprehensive improvement not only expands the scope of application of large language models, but also improves their ability to handle fuzzy or ambiguous information, laying a solid foundation for achieving high-quality human-computer interaction.

[0142] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A text processing method based on a large language model, characterized by: include, Collect text data and generate variant samples through data augmentation technology to obtain a pre-training dataset, introduce the knowledge graph, and output the trained large language model; Adopting an adaptive task selector, the pre-training dataset is automatically matched to the downstream task type. At the same time, an incremental learning framework and meta-learning algorithm are used to optimize hyperparameters. Consider non-textual data, use multimodal models to capture different types of contextual clues, and utilize cross-cultural cognitive frameworks to understand expressions; Based on the updated large language model, a confidence assessment method is established and high-confidence text processing results are output; Based on high-confidence text processing results, build personalized user portraits and optimize recommended content; Collect text data and generate variant samples through data enhancement technology to obtain pre-training data sets, introduce knowledge graphs, and output the trained large language model, including the following steps: Determine the target area and identify the main information sources within the target area; Use crawler technology to access selected primary information sources and download text data; Use synonym replacement, sentence reorganization and context expansion methods to obtain variant samples; Integrate text data and variant samples to generate pre-training datasets; Analyze the unique expressions, metaphors, and slang of various cultures to form a cultural feature tag library, and use bilingual comparison to build a parallel corpus; Apply the BERT-based NER model to automatically identify entities in text and label entity types; Through dependency parsing, the grammatical structure of the sentence is analyzed to find the potential relationship between entities; Use the TransE algorithm to optimize triple representation and map entities and relations into a low-dimensional vector space; When entity and relationship capture fails, re-collect and analyze text data; When entities and relationships are captured successfully, a knowledge graph is constructed based on the captured entities and relationships, and a large language model is trained to obtain a trained large language model. Adopting an adaptive task selector, the pre-training dataset is automatically matched to the downstream task type. At the same time, the incremental learning framework and meta-learning algorithm are used to optimize hyperparameters, including the following steps: Define downstream task types and define feature sets and evaluation metrics for each task type; Use a meta-learning algorithm to train an adaptive task selector based on the defined downstream task type, input a pre-training dataset, and output the predicted downstream task type and hyperparameter configuration recommendations; Regularly obtain the latest text data from information sources, clean and annotate it, and add it to the parallel corpus. Then merge it with historical text data to form a new pre-training dataset. Use the mini-batch gradient descent optimization method combined with the prediction results to update some hyperparameters of the trained large language model. Considering non-text data, using multimodal models to capture different types of context clues, and using a cross-cultural cognitive framework to understand expressions, includes the following steps: Collect non-textual data related to the target domain; Using a multimodal deep learning architecture, text data and non-text data are projected into the same dimension and jointly represented in a unified vector space to form multimodal data; Pair multimodal data with text data to form joint training samples, and use an end-to-end approach to simultaneously train the text encoder and the encoder of other modalities; Using conditional random field technology, we add a cross-cultural adaptation layer to the large language model to perform cultural sensitivity analysis on the input text and identify the cultural characteristics involved; Use parallel corpora from multiple languages ​​and cultures for pre-training, so that the large language model can initially grasp the expressions of different cultures and obtain an updated large language model.

2. The text processing method based on a large language model according to claim 1, wherein: The text data includes authoritative documents, professional forums and academic journals; The non-text data includes images and tables.

3. The text processing method based on a large language model according to claim 1, wherein: Based on the updated large language model, a confidence assessment method is established and high-confidence text processing results are output, including the following steps: Define confidence metrics; The confidence index is Z-Score standardized to a numerical value, and the confidence score is calculated based on the confidence index, which is expressed as ; in, represents the confidence score, Indicates the The dynamic weight of the confidence indicator, Indicates the The standardized score of the confidence index, represents the index of the confidence indicator, Indicates the number of confidence indicators; Set confidence thresholds based on historical text data; When the confidence score is less than or equal to the confidence threshold, the prediction result is judged to be low confidence and marked for further review; When the confidence score is greater than the confidence threshold, the prediction result is judged to be high confidence; Arrange all high-confidence prediction results in descending order and output the text processing result with the highest confidence score.

4. The text processing method based on a large language model according to claim 1, wherein: Based on the high-confidence text processing results, build personalized user portraits and optimize recommended content, including the following steps: Collect user interaction data from multiple channels to form user activity logs; Use natural language processing technology to extract personalized features from user activity logs; Build multi-dimensional user portraits based on the extracted personalized features; Based on high-confidence text processing results, mark interest tags in user portraits; Provide users with customized recommended content based on the interest tags in their user portraits.

5. A text processing system based on a large language model, based on the text processing method based on a large language model according to any one of claims 1 to 4, characterized in that: include, The pre-training module is responsible for collecting text data and generating variant samples through data augmentation technology to obtain a pre-training dataset, introduce the knowledge graph, and output the trained large language model; The fine-tuning module is responsible for automatically matching the pre-training dataset to the downstream task type using an adaptive task selector. At the same time, it uses an incremental learning framework and meta-learning algorithm to optimize hyperparameters. The fusion module is responsible for considering non-text data, using multimodal models to capture different types of contextual clues, and using a cross-cultural cognitive framework to understand expressions; The evaluation module is responsible for establishing a confidence assessment method based on the updated large language model and outputting high-confidence text processing results; The personalization module is responsible for building personalized user portraits and optimizing recommended content based on high-confidence text processing results.

6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the text processing method based on a large language model according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the text processing method based on a large language model according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Text translation model training method and device and text translation method and device

    CN117094418A

  • Question and answer processing method and device, storage medium and program product

    CN118939780A