Text processing method and device, electronic equipment and storage medium

By using pre-trained meta-learning models to generate embedded representations of unlearned texts and adding them to the embedding matrix of large language models, the challenges of large language models in the comprehension and adaptation of new conceptual texts are solved, and the generalization ability and adaptability of the model are improved.

CN120181091APending Publication Date: 2025-06-20AVATR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411315438.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Large language models face challenges in learning and adapting to unlearned texts, especially in the understanding and adaptation of new conceptual texts.

Method used

Generate embedded representations of unlearned texts through pre-trained meta-learning models and add them to the embedding matrix of large language models, enabling large language models to understand the semantics of unlearned texts based on these embedded representations.

Benefits of technology

The ability of large language models to quickly understand and adapt to unlearned texts is realized, the generalization ability and adaptability of the model is improved, and overfitting of new task data is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181091A_ABST
    Figure CN120181091A_ABST
Patent Text Reader

Abstract

The invention discloses a text processing method and device, electronic equipment and a storage medium. The method comprises the steps that a first embedded representation corresponding to a first text to be processed is determined, and the first embedded representation comprises a second embedded representation corresponding to a second text which is not learned by a first language model, the second embedded representation is generated by a pre-trained meta-learning model based on at least one third text containing the second text; and inputting the first embedded representation into the first language model to obtain a processing result output by the first language model. According to the method, the first language model can quickly understand the semantics of the unlearned second text, and the generalization ability and adaptability of the first language model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence (AI) technology, and in particular, but not limited to, a text processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence and natural language processing technologies, large language models (LLMs) have become the core components of various applications. These models learn rich language knowledge and context understanding capabilities through pre-training on large-scale text data.

[0003] However, although large language models perform well in many text processing tasks, they still face challenges in quickly learning and adapting to unseen text (such as new concept text). Summary of the Invention

[0004] Embodiments of this application provide a text processing method, apparatus, electronic device, and storage medium, enabling large language models to quickly understand and adapt to unseen text.

[0005] The technical solution of this application is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a text processing method, including:

[0007] Determine a first embedding representation corresponding to a first text to be processed, where the first embedding representation includes a second embedding representation corresponding to a second text not learned by a first language model, and the second embedding representation is generated by a pre-trained meta-learning model based on at least one third text including the second text;

[0008] Input the first embedding representation into the first language model to obtain a processing result output by the first language model.

[0009] In some embodiments, before determining the first embedding representation corresponding to the first text to be processed, the method further includes:

[0010] Input at least one third text including the second text into the meta-learning model to obtain at least one context embedding vector output by the meta-learning model, and the at least one context embedding vector corresponds to the at least one third text one by one;

[0011] Determine the second embedding representation based on the at least one context embedding vector.

[0012] In some embodiments, determining the second embedding representation based on the at least one context embedding vector includes:

[0013] Aggregate at least one context embedding vector to obtain an aggregate vector;

[0014] Perform a linear transformation on the aggregate vector to obtain a second embedding representation.

[0015] In some embodiments, aggregating at least one context embedding vector to obtain an aggregate vector includes:

[0016] Based on the average pooling algorithm, aggregate at least one context embedding vector to obtain an aggregate vector.

[0017] In some embodiments, after determining a second embedding representation based on at least one context embedding vector, the method further includes:

[0018] Add the second embedding representation to the embedding matrix of the first language model so that the first language model can understand the semantics of the second text based on the second embedding representation.

[0019] In some embodiments, the meta-learning model is constructed based on a second language model with frozen parameters.

[0020] In some embodiments, the second language model is a masked language model.

[0021] In a second aspect, an embodiment of the present application provides a text processing apparatus, including:

[0022] A determination module, configured to determine a first embedding representation corresponding to a first text to be processed, where the first embedding representation includes a second embedding representation corresponding to a second text not learned by the first language model, and the second embedding representation is generated by a pre-trained meta-learning model based on at least one third text including the second text;

[0023] An obtaining module, configured to input the first embedding representation into the first language model to obtain a processing result output by the first language model.

[0024] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, the text processing method provided in the first aspect is implemented.

[0025] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is run by a processor, the text processing method provided in the first aspect is implemented.

[0026] Fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the text processing method provided in the first aspect above.

[0027] The beneficial effects of the technical solution provided by the embodiment of the present application compared with the prior art are as follows:

[0028] In the present application, when the first text to be processed includes a second text that the first language model has not learned, the pre-trained meta-learning model can generate a second embedding representation corresponding to the second text based on at least one third text including the second text. At this time, the first language model can obtain a first embedding representation corresponding to the first text and including the second embedding representation based on the second embedding representation. Then, the first language model can process the first embedding representation including the second embedding representation to obtain a corresponding processing result. In this process, the first language model can quickly understand and adapt to the unlearned second text, effectively improving the generalization ability and adaptability of the first language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic structural diagram of a text processing system provided by an embodiment of the present application;

[0030] Figure 2 It is one of the schematic flowcharts of a text processing method provided by an embodiment of the present application;

[0031] Figure 3 It is another schematic flowchart of a text processing method provided by an embodiment of the present application;

[0032] Figure 4 It is yet another schematic flowchart of a text processing method provided by an embodiment of the present application;

[0033] Figure 5 It is a schematic diagram of a text processing principle provided by an embodiment of the present application;

[0034] Figure 6 It is still another schematic flowchart of a text processing method provided by an embodiment of the present application;

[0035] Figure 7 It is a schematic diagram of the composition structure of a text processing device provided by an embodiment of the present application;

[0036] Figure 8 It is a schematic diagram of the hardware entity of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will further describe the specific technical solutions of the application in detail in conjunction with the accompanying drawings in the embodiments of this application. The following embodiments are used to illustrate this application, but are not used to limit the scope of this application.

[0038] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0039] In the following description, the terms "first / second / third" are only used to distinguish different objects, and do not represent a specific order for the objects, and there is no limitation on the order of precedence. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0041] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this invention are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0042] 1) Embedding representation: Refers to a feature vector or feature matrix used to represent information such as text, audio, and video. Mathematically, it represents a mapping relationship. For example, F: X -> Y. Where F represents the mapping, and X and Y represent any two non-empty sets, and X and Y have a mapping relationship. In natural language processing, the embedding representation specifically refers to the mapping result from the semantic space to the vector space. For example, a word can be represented by an embedding representation in the form of a low-dimensional vector. Then, for this word, the mapping relationship from the semantic space to the vector space can be expressed as F: X -> Y, where F represents the mapping relationship, X is the word, and Y is the embedding representation. Both X and Y are non-empty sets.

[0043] 2) Large language model: Refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks, such as text classification, question answering, and dialogue, etc., and is an important approach to artificial intelligence.

[0044] 3) Fine-tuning: Based on a general large model, for areas beyond the scope or not proficient in, use specialized datasets or methods to make corresponding adjustments and optimizations to improve its applicability and completion in specific fields or tasks.

[0045] In related technologies, large language models can rely on the fine-tuning process to learn new concept texts. This process requires a large amount of labeled data and computing resources to fine-tune the model's parameters, which is not conducive to quickly adapting to the emergence of new concepts.

[0046] In some embodiments, large language models can match context prompting methods to learn new concept texts. However, although the context prompting method can guide the model to focus on new concepts, this method lacks robustness, is easily interfered by the context, and cannot effectively transmit detailed information about new concepts.

[0047] In some embodiments, large language models can splice the embedding representations obtained by traditional few-shot learning methods to learn new concept texts. In this process, the embedding representations generated by traditional few-shot learning methods (such as methods based on global word vectors) are difficult to be compatible with the complex representation space of large language models and are not applicable to large language models.

[0048] To solve at least some of the above technical problems, embodiments of the present application provide a text processing method, device, electronic device, and storage medium, so that large language models can quickly understand and adapt to unlearned texts.

[0049] It should be noted that the text processing method provided by the embodiments of the present application can be applied to a text processing system. Figure 1 The following is a schematic diagram of the architecture of a text processing system provided by the embodiments of the present application. Refer to Figure 1 , to implement the processing application that supports the text processing method, the terminal devices (terminal device 400-1 and terminal device 400-2 are shown in the figure) are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0050] In some embodiments, the terminal device sends the first text input by the user to the server 200. The server 200 obtains the first text to be processed, where the first text includes a second text, and the second text is a text that the first language model has not learned. The server 200 inputs the first embedding representation corresponding to the first text into the first language model to obtain the processing result output by the first language model.

[0051] In some embodiments, the server 200 sends the processing result output by the first language model to the terminal device, and the terminal device displays the processing result output by the first language model on the graphical interface (graphical interfaces 410-1 and 410-2 are shown in the figure).

[0052] In some embodiments, the terminal device may be a vehicle-mounted device, and the vehicle-mounted device is installed in the vehicle equipment.

[0053] In some embodiments, an AI processor may be installed on the terminal device, and the AI processor may execute the steps of text processing performed by the above-mentioned server 200. In this case, the terminal device may not communicate with the server 200.

[0054] An exemplary introduction to a text processing method provided by an embodiment of the present application is given below.

[0055] Figure 2 FIG. is one of the flow diagrams of a text processing method provided by an embodiment of the present application. Refer to Figure 2 As shown, taking the server as the execution subject, the text processing method may include steps S201 to S202.

[0056] Step S201: Determine a first embedding representation corresponding to a first text to be processed, where the first embedding representation includes a second embedding representation corresponding to a second text not learned by the first language model, and the second embedding representation is generated by a pre-trained meta-learning model based on at least one third text including the second text.

[0057] In some embodiments, the first text to be processed may be obtained first, and then the first embedding representation corresponding to the first text to be processed may be determined. If the first text includes a second text not learned by the first language model, the second embedding representation corresponding to the second text may be generated by a pre-trained meta-learning model based on at least one third text including the second text, and the embedding representations of the remaining texts in the first text except the second text may be obtained based on the embedding layer matrix of the first language model. Furthermore, the second embedding representation corresponding to the second text and the embedding representations of the remaining texts in the first text except the second text may be combined to obtain the first embedding representation corresponding to the first text.

[0058] It should be noted that the at least one third text including the second text may be at least one phrase, sentence, paragraph, etc., and the embodiments of the present application do not make specific limitations thereto.

[0059] It can be understood that the server can obtain the first text to be processed through the terminal device. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, and big data and artificial intelligence platforms.

[0060] In some embodiments, the first text can be the text input by the user into the terminal device.

[0061] In some embodiments, the first text can be the text input by the user using the input control, or the text corresponding to the voice information input by the user.

[0062] In some embodiments, the first text can be a text containing multiple words or elements such as a sentence, a paragraph, etc. In one embodiment, the first text can be a Chinese text, an English text, or a text corresponding to other languages. The embodiments of the present application do not limit this.

[0063] In some embodiments, the first text may include a second text, and the second text may be a text that the first language model has not learned.

[0064] It can be understood that since the first language model has not learned the second text, the representation space of the first language model cannot accurately represent the second text. That is to say, relative to the first language model, the second text is specific domain data and specific tasks. The first language model does not have the ability to process this specific domain data and specific tasks.

[0065] In some embodiments, the second text can be the text representation corresponding to a new concept; where the new concept can be a new word, a new phrase, a new sentence, or a new paragraph, etc.

[0066] In some embodiments, the first language model can be a large language model. For example, it is a GPT (generative pre-trained) series model. It can be understood that the large language model is pre-trained on a large dataset, shows emerging capabilities, and performs well in various tasks such as language translation, summarization, coding, and question answering.

[0067] In some embodiments, the user requests the processing application on the terminal device to perform text processing and inputs the first text to the processing application. After obtaining the first text, the processing application can send it to the terminal device. After obtaining the first text, the terminal device can send the first text to the server through the network so that the server processes the first text.

[0068] It should be noted that the Meta-Learning Model is a new type of machine learning model. The core idea of meta-learning is to enable the machine learning model to learn on multiple tasks, so as to master a general learning strategy or meta-knowledge. This meta-knowledge can help the model adapt and learn faster on new and unseen tasks, that is, under limited sample data, the model can quickly adapt to new tasks. In the embodiments of the present application, by pre-training or multi-task learning on the meta-learning model, it masters a general learning strategy, extracts meta-knowledge useful for learning tasks, and then for a new task, such as extracting a second embedding representation corresponding to a second text, it only needs a small number of samples (i.e., at least one third text), and based on the learning strategy it has pre-mastered, it can quickly extract the second embedding representation corresponding to the second text.

[0069] In some embodiments, the meta-learning model may include a self-attention layer for extracting the context information of the text, so that the obtained embedding representation corresponding to the text can reflect the context information of the text.

[0070] In some embodiments, during the pre-training process of the meta-learning model, different optimization algorithms can be used to optimize the meta-learning model, such as the Stochastic Gradient Descent (SGD) optimization algorithm, the Adam optimization algorithm, and the optimization algorithm based on the Root Mean Square Propagation (RMSprop), etc., or adjust the hyperparameters of these algorithms to make the meta-learning model have better performance.

[0071] Step S202: Input the first embedding representation into the first language model to obtain the processing result output by the first language model.

[0072] It can be understood that the server can input the first embedding representation corresponding to the first text into the first language model, so that the first language model processes or infers and analyzes the first embedding representation to obtain the corresponding processing result, and this processing result can be a text representation.

[0073] In some embodiments, the server can input the first embedding representation corresponding to the first text into the network layer of the first language model, so that the network layer of the first language model processes or infers and analyzes the first embedding representation to obtain the corresponding processing result.

[0074] In some embodiments, the network layer of the first language model can be constructed by a neural network architecture Transformer based on an attention mechanism.

[0075] It should be noted that after the first language model processes the first embedding representation, the corresponding output embedding representation is first obtained, and then the corresponding text representation is obtained from the output embedding representation as the processing result.

[0076] It should be noted that the text corresponding to the output embedding representation obtained after the first language model processes the first embedding representation may also include other texts that the first language model has not learned, and the output embedding representations corresponding to these other texts that have not been learned can also be generated by a pre-trained meta-learning model based on at least one text containing these other texts that have not been learned.

[0077] In some embodiments, after the server obtains the processing result output by the first language model, it can send the processing result to the terminal device through the network. After obtaining the processing result, the terminal device can send it to the processing application. After obtaining the processing result, the processing application can display it to the user, thereby completing text processing.

[0078] It should be noted that the processing result output by the first language model can be a text obtained by completing the first text, or a text obtained by answering the first text.

[0079] It can be understood that in the text processing method provided by the embodiments of the present application, when the second text that the first language model has not learned is included in the first text to be processed, the pre-trained meta-learning model can generate the second embedding representation corresponding to the second text based on at least one third text containing the second text. At this time, the first language model can obtain the first embedding representation containing the second embedding representation corresponding to the first text based on the second embedding representation, and then the first language model can process the first embedding representation containing the second embedding representation to obtain the corresponding processing result. In this process, the first language model can quickly understand and adapt to the unlearned second text, effectively improving the generalization ability and adaptability of the first language model.

[0080] In some possible implementation manners, before determining the first embedding representation corresponding to the first text to be processed, the method further includes:

[0081] Inputting at least one third text containing the second text into the meta-learning model to obtain at least one context embedding vector output by the meta-learning model, where the at least one context embedding vector corresponds to the at least one third text one by one;

[0082] Determining the second embedding representation based on the at least one context embedding vector.

[0083] In some embodiments, the meta-learning model may include a self-attention layer, which can extract the context embedding vectors of the input text. Among them, the self-attention layer may be constructed by a neural network architecture Transformer based on the attention mechanism.

[0084] In some embodiments, the self-attention layer may be a single-head self-attention layer, a multi-head self-attention layer, a self-attention layer with relative position information, etc., and the embodiments of the present application do not make specific limitations thereto.

[0085] In the embodiments of the present application, at least one third text containing a second text may be input into the meta-learning model including a self-attention layer to obtain at least one context embedding vector output by the meta-learning model, where at least one context embedding vector corresponds to at least one third text one by one, and then based on at least one context embedding vector, a second embedding representation corresponding to the second text is determined.

[0086] Exemplarily, Figure 3 is the second schematic diagram of the process of a text processing method provided by the embodiments of the present application. As Figure 3 shown, the method includes:

[0087] S301. Input at least one third text containing a second text into the meta-learning model to obtain at least one context embedding vector output by the meta-learning model, and the at least one context embedding vector corresponds to the at least one third text one by one;

[0088] S302. Determine a second embedding representation based on the at least one context embedding vector;

[0089] S303. Determine a first embedding representation corresponding to the first text to be processed, and the first embedding representation includes the second embedding representation corresponding to the second text not learned by the first language model;

[0090] S304. Input the first embedding representation into the first language model to obtain a processing result output by the first language model.

[0091] It can be understood that in the embodiments of the present application, at least one context embedding vector corresponding to at least one third text one by one can reflect the context information of the second text in each third text, and then a second embedding representation corresponding to the second text is obtained from the at least one context embedding vector, which can make the second embedding representation more accurately represent the semantics of the second text, and further enable the first language model to more accurately understand the semantics of the second text.

[0092] In some possible implementation manners, the determining the second embedding representation based on the at least one context embedding vector includes:

[0093] Perform an aggregation process on the at least one context embedding vector to obtain an aggregated vector;

[0094] Perform a linear transformation on the aggregated vector to obtain the second embedding representation.

[0095] In the embodiments of the present application, after obtaining at least one context embedding vector output by the meta-learning model, an aggregation process can be performed on the at least one context embedding vector, that is, the at least one context embedding vector can be fused into one vector, which is the aggregated vector, and then a linear transformation is performed on the aggregated vector to obtain the second embedding representation corresponding to the second text.

[0096] In some embodiments, an aggregation process can be performed on the at least one context embedding vector based on any one of aggregation methods such as average pooling, max pooling, weighted summation, and attention mechanism to obtain an aggregated vector.

[0097] It should be noted that performing a linear transformation on the aggregated vector can change the dimension of the aggregated vector, converting the aggregated vector from a high-dimensional feature space to a low-dimensional feature space, or from a low-dimensional feature space to a high-dimensional feature space, so that the dimension of the obtained second embedding representation can match the dimension of the vector representation that the first language model can process.

[0098] It should be noted that since different large language models have different internal structures and numbers of parameters, this directly affects the dimensions of the vector representations they can process. Therefore, in the embodiments of the present application, by performing a linear transformation on the aggregated vector, the dimension of the second embedding representation can be made to conform to the dimension of the vector representation that the first language model can process.

[0099] Exemplarily, Figure 4 is the third flowchart diagram of a text processing method provided by the embodiments of the present application. As Figure 4 shown, the method includes:

[0100] S401. Input at least one third text containing the second text into the meta-learning model to obtain at least one context embedding vector output by the meta-learning model, where the at least one context embedding vector corresponds one-to-one to the at least one third text;

[0101] S402. Perform an aggregation process on the at least one context embedding vector to obtain an aggregated vector;

[0102] S403. Perform a linear transformation on the aggregated vector to obtain the second embedding representation;

[0103] S404. Determine a first embedding representation corresponding to a first text to be processed, where the first embedding representation includes the second embedding representation corresponding to a second text that has not been learned by the first language model;

[0104] S405. Input the first embedding representation into the first language model to obtain a processing result output by the first language model.

[0105] It can be understood that in the embodiments of the present application, after obtaining at least one context embedding vector output by the meta-learning model, the at least one context embedding vector is aggregated to obtain an aggregated vector, and then a linear transformation is performed on the aggregated vector, so that the dimension of the second embedding representation corresponding to the second text finally obtained can conform to the dimension of the vector representation that the first language model can process, which helps the first language model to understand the semantics of the second text based on the second embedding representation to complete the corresponding task.

[0106] In some possible implementation manners, the aggregating the at least one context embedding vector to obtain an aggregated vector includes:

[0107] Based on the average pooling algorithm, aggregate the at least one context embedding vector to obtain the aggregated vector.

[0108] In the embodiments of the present application, the at least one context embedding vector can be aggregated based on the average pooling algorithm to obtain an aggregated vector.

[0109] It should be noted that the average pooling algorithm can be used for feature dimensionality reduction. It performs an average operation on a part of the elements in the vector (i.e., a sub-vector or the elements within a window), and then uses the average value as the output of this region, avoiding deviations caused by individual extreme values, and can effectively help the model capture background information or global statistical features.

[0110] It can be understood that in the embodiments of the present application, by aggregating the at least one context embedding vector based on the average pooling algorithm, the obtained aggregated vector can more comprehensively and accurately represent the semantics of the second text, which helps the subsequent first language model to accurately understand the semantics of the second text.

[0111] In some possible implementation manners, after determining the second embedding representation based on the at least one context embedding vector, the method further includes:

[0112] Add the second embedding representation to the embedding matrix of the first language model so that the first language model can understand the semantics of the second text based on the second embedding representation.

[0113] In the embodiments of the present application, after obtaining the second embedding representation corresponding to the second text, the second embedding representation can be added to the embedding matrix of the first language model, so that the first language model can understand the semantics of the second text based on the second embedding representation.

[0114] It can be understood that in the embodiments of the present application, by adding the second embedding representation to the embedding matrix of the first language model, the scale expansion of the embedding matrix of the first language model is realized.

[0115] It can be understood that after adding the second embedding representation to the embedding matrix of the first language model, the first embedding representation corresponding to the first text containing the second text can be directly obtained by using the first language model. Furthermore, the first language model processes the first embedding representation to obtain the final processing result.

[0116] It should be noted that each large language model is usually equipped with an embedding matrix, which is used to convert text data into vector representations, enabling the model to better understand the text content. The embedding matrix assigns a unique index to each word, and each row of this matrix corresponds to the vector representation of a word, which can capture the semantic information of the word. By looking up the row vector corresponding to each word, the original text data is converted into the form of vector representations. In this way, each sentence is represented as a series of vectors, and these vectors retain the semantic information in the original text.

[0117] It should be noted that in the input layer of the large language model, the embedding matrix is used to convert the input data (such as text) into vector representations so that the model can process it. This conversion helps the model better understand and process the input data. In the output layer, the embedding matrix can convert the output of the model back into a human - understandable format or provide an appropriate representation for the next - step processing. Therefore, the input layer and output layer of the large language model respectively include corresponding embedding matrices.

[0118] In some embodiments, after aggregating at least one context embedding vector to obtain an aggregated vector, the aggregated vector can be subjected to a first linear transformation, and then the embedding representation obtained after the first linear transformation is added to the embedding matrix of the input layer of the first language model; and, the aggregated vector can be subjected to a second linear transformation, and then the embedding representation obtained after the second linear transformation is added to the embedding matrix of the output layer of the first language model. Wherein, the linear transformation algorithms corresponding to the first linear transformation and the second linear transformation can be determined based on the internal structure and / or internal parameters of the first language model.

[0119] It can be understood that in the embodiments of the present application, by obtaining the second embedding representation corresponding to the second text not learned by the first language model based on the meta-learning model, and then adding the second embedding representation to the embedding matrix of the first language model, the first language model can quickly understand the semantics of the second text not learned, and improve the generalization ability and adaptability of the first language model.

[0120] In some possible implementation manners, the meta-learning model is constructed based on a second language model with frozen parameters.

[0121] It should be noted that a model with frozen parameters refers to that during the model training process, some parameters in the model are not updated (i.e., "frozen") and remain unchanged. If some layers of the pre-trained model have well captured useful features and it is not desired to destroy these features during the fine-tuning process, the parameters of these layers can be frozen, which can reduce the number of parameters to be optimized and speed up the training speed. When the target task is similar but not exactly the same as the pre-trained task, most layers of the pre-trained model can be frozen, and only the last layer or several layers are fine-tuned to adapt to the new task, which helps to utilize the general features learned by the pre-trained model on a large amount of data while reducing overfitting to the new task data.

[0122] In the embodiments of the present application, a meta-learning model can be constructed based on a second language model with frozen parameters to reduce the number of parameters to be optimized by the meta-learning model and reduce overfitting of the meta-learning model to the new task data. In some embodiments, the second language model can be a pre-trained language model.

[0123] It should be noted that pre-trained language models (PLMs) can capture rich representations and context information of language through unsupervised learning on large-scale text data. These models can be fine-tuned on a variety of natural language processing tasks to meet the requirements of specific tasks without having to train the model from scratch.

[0124] In some embodiments, the second language model can be any one of models such as BERT (Bidirectional Encoder Representations from Transformers), an improved version of the BERT model RoBERTa (Robustly Optimized BERT Pretraining Approach), LLaMA (Large Language Model Meta AI), and GPT (Generative Pre-trained Transformer) model.

[0125] It can be understood that by constructing a meta-learning model based on a second language model with parameter freezing, the embodiments of the present application can not only reduce the number of parameters to be optimized in the meta-learning model, but also reduce the overfitting of the meta-learning model to new task data.

[0126] In some possible implementation manners, the second language model is a masked language model (MLM).

[0127] It should be noted that MLM is a self-supervised learning technique. Its core idea is that in the pre-training stage of the model, by randomly masking some words in the input text (or replacing them with special tokens such as <mask>) and requires the model to predict these masked words based on the context, thereby learning rich language representations. This training method enables the model to deeply understand the context of words and their relationships with other words in the sentence, thereby enhancing the model's language understanding and generation capabilities.

[0128] In the embodiments of the present application, a meta-learning model can be constructed based on a masked language model with parameter freezing, enabling the meta-learning model to deeply understand the context of words and their relationships with other words in the sentence, thereby enhancing the model's language understanding and generation capabilities.

[0129] Next, a specific example is used to illustrate the text processing method provided by the embodiments of the present application.

[0130] Figure 5 It is a schematic diagram of the text processing principle provided by the embodiments of the present application. Refer to Figure 5 During the text processing, a large language model (i.e., the first language model in the foregoing embodiments) and a meta-learning model are involved.

[0131] In some embodiments, Figure 5 The large language model shown may include an input embedding layer 501, a network layer 502, and an output embedding layer 503. Among them, the input embedding layer 501 is connected to the linear layer 504 of the meta-learning model, and the output embedding layer 503 is connected to the linear layer 505 of the meta-learning model.

[0132] In some embodiments, in addition to the linear layer 504 and the linear layer 505, Figure 5 The meta-learning model shown may further include a self-attention layer 506, a masked network layer 507, and a masked processing layer 508. Among them, the masked network layer 507 and the masked processing layer 508 can form the masked language model of the meta-learning model.

[0133] In some embodiments, Figure 5 The process of the meta - learning model for text processing shown can include: inputting the support sequence of the second text (i.e., at least one third text containing the second text in the foregoing embodiments) into the masking layer 508 of the meta - learning model. The masking layer 508 can perform masking processing on the second text within the third text and sequentially input the masked text into the masking network layer 507 and the self - attention layer 506 of the meta - learning model for feature extraction, obtaining the embedding representations 509 corresponding to each word in each third text. Then, the meta - learning model performs an aggregation process on the embedding representations corresponding to each word in each third text to obtain the embedding representations 510 corresponding to each third text. Further, an aggregation process is performed on the embedding representations 510 corresponding to each third text to obtain the embedding representation 511. Finally, the embedding representation 511 undergoes linear transformation processing through the linear layer 504 and the linear layer 505, obtaining two embedding representations corresponding to the second text and adding them to the embedding matrix of the input embedding layer 501 and the embedding matrix of the output embedding layer 503 of the large - language model respectively.

[0134] In some embodiments, Figure 5 The process of the large - language model for text processing shown can include: the first text 500 to be processed is input into the input embedding layer 501 of the large - language model, generating the embedding representation of the first text 500. The embedding representation of the first text is input into the network layer 502 of the first language model. The network layer 502 understands the first text based on the embedding representation of the first text, obtaining the output embedding representation, and this output embedding representation is transformed through the output embedding layer 503 of the first language model to obtain the target text 512. It should be noted that the first text contains a second text that the large - language model has not learned. Since the second embedding representation corresponding to the second text obtained by the meta - learning model is pre - added to the input embedding layer 501 and the output embedding layer 503 of the large - language model, the large - language model can quickly understand the second text that it has not learned.

[0135] Figure 6 This is the fourth flowchart of a text processing method provided by an embodiment of the present application. Combining Figure 5 with the text processing principle schematic diagram shown, the text processing method can include steps S601 to S605.

[0136] Step S601: Input the support sequence containing the new concept text into the masking layer of the meta - learning model, so that the masking layer performs masking and marking processing on the new concept text in the support sequence.

[0137] It should be noted that in the embodiments of the present application, the second text is used as the new concept text, and multiple third texts form a support sequence for illustration. In one embodiment, the new concept text may be a new-born text with an unknown meaning. For example, the new concept text may be beije_flag. In beije_flag, beije is the known person name Beyer, flag is the known word flag, and beije_flag is a new-born text with an unknown meaning.

[0138] It can be understood that first, the few-shot support sequences (sentences containing the new concept text) can be input into the masking layer of the meta-learning model for masking. Among them, the new concept text will be replaced with a mask token, for example <mask>Marker.

[0139] It can be understood that after masking the new concept text, the meta-learning model can focus on learning these new concepts.

[0140] In one example, the few-shot support sequences can be K support sequences {s1,..., s K}, where s1 to s K respectively represent the first support sequence to the Kth support sequence, and K is a positive integer. After the K support sequences are input into the masking processing layer, K masked sequences containing masked markers are generated <mask>Support sequence.

[0141] In one example, such as Figure 5 shown, the new concept text beije_flag can be replaced with a masked token <mask>。

[0142] Step S602: Based on the masked network layer and self-attention layer of the meta-learning model, obtain the embedding representations corresponding to each support sequence.

[0143] It can be understood that the masked network layer of the meta-learning model, such as the RoBERTa model, can tokenize and encode K support sequences to extract features and obtain the initial context embedding vectors of the K support sequences. Then, the initial context embedding vectors can be passed through the self-attention layer of the meta-learning model to obtain the embedding representation {h i} of each word in each support sequence, where h i represents the embedding representation of each word in each support sequence.

[0144] In some embodiments, the embedding representations of each word in each support sequence can be aggregated according to formula (1) to obtain the embedding representations {e1,..., e k} corresponding to each of the K support sequences. Among them, e1 to e k respectively represent the embedding representation corresponding to the first support sequence to the embedding representation corresponding to the Kth support sequence in the support sequence. Formula (1) can be:

[0145]

[0146] Among them, n represents that each support sequence has n words, and n is a positive integer; represents the cumulative sum of the embedding representations of the words from the first word to the last word in each support sequence; e i represents the embedding representation corresponding to the ith support sequence.

[0147] It can be understood that the masked network layer and self-attention layer of the meta-learning model are designed to capture semantic information and context information in the support sequence.

[0148] Step S603: Aggregate the embedding representations corresponding to each support sequence to obtain an aggregated vector.

[0149] In some embodiments, the embedding representations corresponding to each support sequence can be aggregated according to formula (2) to obtain an aggregated vector e new (i.e., the embedding representation of the new concept text). Formula (2) can be:

[0150]

[0151] Step S604: Perform a linear transformation on the aggregated vector based on the linear layer of the meta-learning model to obtain the embedding representation of the new concept text.

[0152] It should be noted that in order to make the dimensionality of the embedded representation of the new concept text conform to the vector dimensionality that the large language model can process, two different linear layers based on the meta-learning model can be used to perform linear transformations on e new respectively to generate the embedded representation e in and e out of the new concept text, that is, [ein] = Linear1(enew), [e out = Linear2(e new ), where Linear1 and Linear2 represent two different linear transformation functions.

[0153] Step S605: Add the embedded representation of the new concept text to the embedding matrix of the large language model so that the large language model can quickly understand the new concept text when processing tasks.

[0154] In some embodiments, the embedded representations e in and e out of the new concept text can be added to the embedding matrix of the input layer and the embedding matrix of the output layer of the large language model respectively, so that when the large language model is applied to process various tasks (such as fill-in-the-blank questions, definition inference, and question-and-answer tasks), it can quickly understand the new concept text. It should be noted that the addition in the embodiments of the present application can be understood as extending and splicing the embedded representation of the new concept text into the original embedding matrix of the large language model, or integrating the embedded representation of the new concept text into the original embedding matrix of the large language model.

[0155] In the embodiments of the present application, generating the embedded representation of the new concept can help the large language model quickly understand and adapt to the new concept, improve learning efficiency, reduce training costs, and provide important support for the further development of the model.

[0156] Compared with the prior art, the text processing method provided by the embodiments of the present application can at least bring the following beneficial effects:

[0157] 1. Quick learning of new concepts: Integrating the embedded representation of the new concept text into the large language model enables the large language model to quickly understand and adapt to the new concept without the need for cumbersome fine-tuning or specific task training, thus saving time and costs.

[0158] 2. Few-shot learning: Only need to learn a small number of support sequences (such as a small number of example sentences or definitions) to obtain the embedded representation of the new concept text.

[0159] 3. Rich information acquisition: By generating flexible embedded representations, the meta-learning model can provide richer information to support the learning of new concepts, avoid the influence of context interference, and improve the learning effect.

[0160] 4. Zero-shot learning: After the large language model integrates the embedding representation of new concept texts, it can achieve zero-shot learning when learning new concepts without additional task-specific training, that is, complete the learning task of new concepts without additional training.

[0161] 5. Reducing training costs: Compared with traditional fine-tuning methods, the embodiments of the present application use a meta-learning model to achieve a simple and efficient way of generating embedding representations, avoiding the need for a large amount of data and computing resources, and reducing training costs.

[0162] 6. Applicable to multiple tasks: The new concept embeddings generated by the meta-learning model in the embodiments of the present application can be applied to various learning tasks, such as fill-in-the-blank questions and definition inference, etc., improving the generality and applicability of the text processing method.

[0163] In summary, the embodiments of the present application have obvious advantages in reducing learning costs, improving learning efficiency, enhancing generalization ability, simplifying the application process, supporting zero-shot learning, and improving interpretability, providing an effective technical solution for the continuous learning and knowledge expansion of large language models.

[0164] Figure 7 FIG. is a schematic structural diagram of a text processing device provided by an embodiment of the present application. As Figure 7 shown, the text processing device 700 includes:

[0165] A determination module 701, configured to determine a first embedding representation corresponding to a first text to be processed, where the first embedding representation includes a second embedding representation corresponding to a second text not learned by the first language model, and the second embedding representation is generated by a pre-trained meta-learning model based on at least one third text including the second text;

[0166] An obtaining module 702, configured to input the first embedding representation into the first language model to obtain a processing result output by the first language model.

[0167] In some embodiments, the determination module 701 includes:

[0168] An obtaining unit, configured to input at least one third text including the second text into the meta-learning model to obtain at least one context embedding vector output by the meta-learning model, and the at least one context embedding vector corresponds to the at least one third text one by one;

[0169] A determination unit, configured to determine the second embedding representation based on the at least one context embedding vector.

[0170] In some embodiments, the determination unit is further configured to:

[0171] Aggregate the at least one context embedding vector to obtain an aggregated vector;

[0172] Perform a linear transformation on the aggregated vector to obtain the second embedding representation.

[0173] In some embodiments, the determining unit is further configured to:

[0174] Based on the average pooling algorithm, aggregate the at least one context embedding vector to obtain the aggregated vector.

[0175] In some embodiments, the determining unit is further configured to:

[0176] After determining the second embedding representation based on the at least one context embedding vector, add the second embedding representation to the embedding matrix of the first language model, so that the first language model can understand the semantics of the second text based on the second embedding representation.

[0177] In some embodiments, the meta-learning model is constructed based on a second language model with frozen parameters.

[0178] In some embodiments, the second language model is a masked language model.

[0179] The description of the above device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0180] It should be noted that in the embodiments of the present application, if the above text processing method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the related technology can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0181] Figure 8 Schematic diagram of the hardware entity of an electronic device provided by an embodiment of the present application, as shown in Figure 8 shown, the hardware entity of the electronic device 800 includes: a processor 801 and a memory 802. Among them, the memory 802 stores a computer program that can run on the processor 801, and when the processor 801 executes the program, it implements the steps in the method of any of the above embodiments.

[0182] The memory 802 stores a computer program that can run on the processor. The memory 802 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or already processed by the processor 801 and each module in the electronic device 800 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory.

[0183] When the processor 801 executes the program, it implements the steps of the text processing method of any of the above. The processor 801 generally controls the overall operation of the electronic device 800.

[0184] In one embodiment, the electronic device 800 may be a device deployed in a vehicle, such as an in-vehicle device.

[0185] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the text processing method of any of the above embodiments.

[0186] It should be noted here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0187] The above-mentioned processor may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device implementing the functions of the above-mentioned processor may also be other devices, which are not specifically limited in the embodiments of the present application.

[0188] The above-mentioned computer storage medium / memory may be a read-only memory, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above-mentioned memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0189] The embodiments of the present application provide a computer program, including computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes to implement some or all of the steps in the above-mentioned method.

[0190] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. This computer program product can be specifically implemented in a way of hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0191] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program and computer program product are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0192] As described above, only the embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.< / mask> < / mask> < / mask> < / mask>

Claims

1. A text processing method, characterized in that: The method comprises: Determine a first embedding representation corresponding to a first text to be processed, wherein the first embedding representation includes a second embedding representation corresponding to a second text that has not been learned by the first language model, and the second embedding representation is generated by a pre-trained meta-learning model based on at least one third text including the second text; The first embedding representation is input into the first language model to obtain a processing result output by the first language model.

2. The method according to claim 1, characterized in that Before determining the first embedded representation corresponding to the first text to be processed, the method further includes: Inputting at least one third text including the second text into the meta-learning model to obtain at least one context embedding vector output by the meta-learning model, wherein the at least one context embedding vector corresponds one-to-one to the at least one third text; Based on the at least one context embedding vector, the second embedding representation is determined.

3. The method according to claim 2, characterized in that The determining the second embedding representation based on the at least one context embedding vector comprises: Performing aggregation processing on the at least one context embedding vector to obtain an aggregated vector; Perform a linear transformation on the aggregated vector to obtain the second embedding representation.

4. The method according to claim 3, characterized in that The performing aggregating processing on the at least one context embedding vector to obtain an aggregated vector includes: Based on an average pooling algorithm, the at least one context embedding vector is aggregated to obtain the aggregated vector.

5. The method according to any one of claims 2 to 4, characterized in that: After determining the second embedding representation based on the at least one context embedding vector, the method further includes: The second embedding representation is added to the embedding matrix of the first language model, so that the first language model can understand the semantics of the second text based on the second embedding representation.

6. The method according to any one of claims 1 to 4, characterized in that: The meta-learning model is constructed based on a parameter-frozen second language model.

7. The method according to claim 6, characterized in that The second language model is a masked language model.

8. A text processing device, characterized in that: The device comprises: A determination module, configured to determine a first embedding representation corresponding to a first text to be processed, wherein the first embedding representation includes a second embedding representation corresponding to a second text that has not been learned by the first language model, and the second embedding representation is generated by a pre-trained meta-learning model based on at least one third text including the second text; An obtaining module is used to input the first embedding representation into the first language model to obtain a processing result output by the first language model.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the text processing method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text processing method according to any one of claims 1 to 7 is implemented.