Training method for text processing model, training apparatus for text processing model, electronic device, program product, and storage medium
By splicing multiple independent sample texts and using mask attention mechanism for prediction processing, the problem of poor generalization and insufficient multi-round dialogue capabilities in the prior art is solved, and higher text processing model accuracy and multi-round dialogue capabilities are achieved.
Patent Information
- Application Number
- PCT/CN2024/116277
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-09-02
- Publication Date
- 2025-05-15
AI Technical Summary
During the training process of the supervised fine-tuning stage, existing large language models do not have correlations between different training texts in a unified batch, resulting in poor generalization of text processing models and insufficient multi-round dialogue capabilities.
By obtaining the sample text collection, multiple independent sample texts are spliced, longer spliced sample texts are generated, and the mask attention mechanism is used to predict the spliced sample features, determine the loss function and update the parameters of the text processing model.
The accuracy and multi-round dialogue capabilities of the trained text processing model are improved, and the model's generalization ability and context analysis ability of language copywriting knowledge rules are enhanced.
Smart Images

Figure CN2024116277_15052025_PF_FP_ABST
Abstract
Description
Text processing model training method, text processing model training device, electronic device, program product and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on the Chinese patent application with application number 202311484312.X and application date of November 8, 2023, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field
[0003] The present application relates to computer and artificial intelligence technologies, and in particular to a text processing model training method, a text processing model training device, an electronic device, a program product, and a storage medium. Background Art
[0004] Artificial Intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0005] In the prior art, the supervised fine-tuning phase of existing large-scale language models typically uses domain-specific models for training. During the training process, different training texts within the same batch may not necessarily belong to relevant contexts, resulting in poor generalization and multi-turn dialogue capabilities of the batch-trained text processing models. Currently, there is no effective way to improve the accuracy of trained text processing models.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a text processing model training method, a text processing model training device, an electronic device and a computer-readable storage medium, and a computer program product, which can improve the accuracy of text prediction by the trained text processing model.
[0008] The technical solution of the embodiment of the present application is implemented as follows:
[0009] An embodiment of the present application provides a method for training a text processing model, the method being executed by an electronic device and comprising:
[0010] Acquire a sample text set, wherein the sample text set includes: a plurality of independent sample texts and sample labels of the plurality of independent sample texts;
[0011] splicing at least two of the plurality of independent sample texts to obtain a plurality of spliced sample texts;
[0012] Based on the multiple splicing sample texts, the initialized text processing model is called to perform feature extraction processing to obtain multiple splicing sample features;
[0013] Based on each of the spliced sample features, a masked attention mechanism is called to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking the features corresponding to each of the independent sample texts in the spliced sample features respectively;
[0014] Determining a loss function based on a difference between the predicted text and the sample label;
[0015] Parameters of the text processing model are updated based on the loss function to obtain the trained text processing model.
[0016] The present invention provides a training device for a text processing model, including:
[0017] A data acquisition module is configured to acquire a sample text set, wherein the sample text set includes: a plurality of independent sample texts and sample labels of the plurality of independent sample texts;
[0018] The data acquisition module is configured to perform splicing processing on at least two of the multiple independent sample texts to obtain multiple spliced sample texts;
[0019] A model training module is configured to call the initialized text processing model to perform feature extraction processing based on the multiple splicing sample texts to obtain multiple splicing sample features;
[0020] The model training module is configured to call a masked attention mechanism based on each of the spliced sample features to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking the features corresponding to each of the independent sample texts in the spliced sample features respectively;
[0021] The model training module is configured to determine a loss function based on the difference between the predicted text and the sample label;
[0022] The model training module is configured to perform parameter update processing on the text processing model based on the loss function to obtain the trained text processing model.
[0023] An embodiment of the present application provides an electronic device, comprising:
[0024] a memory for storing computer-executable instructions;
[0025] The processor is used to implement the training method of the text processing model provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.
[0026] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing the training method of the text processing model provided in the embodiment of the present application when executed by a processor.
[0027] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the training method of the text processing model provided in the embodiment of the present application is implemented.
[0028] The embodiments of the present application have the following beneficial effects:
[0029] During the model training process, short independent sample texts are spliced into longer spliced texts, which improves the richness of training samples and can save the computing resources required to obtain training samples. The use of spliced texts can improve the model's multi-round dialogue capabilities and the accuracy of the trained text processing model; a masked attention mechanism is used for the spliced sample features of the spliced text samples. Since there is no correlation between independent samples, each independent sample text is masked based on the masked attention mechanism, allowing the model to better generalize the language and copywriting knowledge rules, while enabling the text processing model to have better context parsing capabilities, thereby improving the accuracy of the text processing model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG1 is a schematic diagram of an application mode of a training method for a text processing model provided in an embodiment of the present application;
[0031] FIG2 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0032] FIG3A is a schematic diagram of a first flow chart of a text processing model training method provided in an embodiment of the present application;
[0033] FIG3B is a second flow chart of the text processing model training method provided in an embodiment of the present application;
[0034] FIG3C is a third flow chart of the text processing model training method provided in an embodiment of the present application;
[0035] FIG3D is a schematic diagram of a fourth flow chart of the text processing model training method provided in an embodiment of the present application;
[0036] FIG4 is a first structural diagram of a text processing model provided in an embodiment of the present application;
[0037] FIG5 is a second structural diagram of a text processing model provided in an embodiment of the present application;
[0038] FIG6 is a schematic diagram of the structure of the attention matrix provided in an embodiment of the present application;
[0039] FIG7 is a schematic diagram of the structure of the spliced sample text provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0041] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0042] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0043] In this application, the implementation of the data capture technical solution involved (for example: obtaining text data as training data from the Internet). When the above embodiments of this application are applied to specific products or technologies, the relevant data collection (for example: receiving text to be replied uploaded by users), use and processing processes should comply with the requirements of national laws and regulations, and conform to the principles of legality, legitimacy and necessity. It does not involve obtaining data types prohibited or restricted by laws and regulations, and will not hinder the normal operation of the target website.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0045] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0046] 1) Transformer: This is a deep neural network model based on the self-attention mechanism and is widely used in various tasks in natural language processing, such as text classification, machine translation, and question-answering systems. This model can convert an input sequence into an output sequence while retaining important information in the input sequence. Due to its excellent performance in processing long texts, the Transformer model has been widely used in the field of Chinese natural language processing. Compared to traditional recurrent neural networks (RNNs) and convolutional neural networks (CNNs), the Transformer model can perform parallel computations, speeding up training. It has been widely used in tasks such as natural language processing, speech recognition, and image generation.
[0047] 2) Natural Language Processing (NLP): It is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; it also involves important technologies for model training in computer science, mathematics, and artificial intelligence. Pre-trained models are developed from large language models (LLMs) in the field of NLP. After fine-tuning, large language models can be widely used in downstream tasks. Natural language processing technologies generally include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.
[0048] 3) Large Language Model (LLM): A deep learning model trained using large amounts of text data can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question answering, and conversation, and are an important path to artificial intelligence.
[0049] 4) Normalization (Softmax) function: A function used to convert the output values of different categories into a probability distribution in the range [0, 1] and 1. The formula is as follows: Among them, Z i is the output value of the i-th node, and C is the number of output nodes, that is, the number of classification categories.
[0050] 5) Attention Mechanism: The attention mechanism is a neural network structure that enables a model to "focus" on certain specific parts of input information while ignoring others. In fields such as natural language processing and computer vision, the attention mechanism is widely used to understand key information in text or images, improving model performance and accuracy. The attention mechanism in deep learning is essentially similar to the selective visual attention mechanism in humans. Its core goal is to select the information that is most critical to the current task objective from a large amount of information.
[0051] 6) Masking: This is a data processing technique that controls the visibility and usability of data by shielding or hiding certain parts of the data. It is used in natural language processing and machine learning (ML). In natural language processing, masking is used to control a model's access to specific words or characters in a text sequence. For example, masking is used in language models to prevent the model from accessing future input data.
[0052] 7) Masked Attention: An attention mechanism used in natural language processing (NLP). It is applied to tasks such as machine translation, speech recognition, and language modeling. In the masked attention mechanism, the mask is a binary matrix of the same size as the input sequence, where 1s represent locations to be masked and 0s represent accessible locations. The mask is used to block certain parts of the sequence to ensure that the model does not use future information to influence current predictions.
[0053] The embodiments of the present application provide a text processing model training method, a text processing model training device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of text prediction by the trained text processing model.
[0054] The following describes exemplary applications of electronic devices provided by embodiments of the present application. The electronic devices provided by embodiments of the present application can implement terminal devices, such as laptop computers, tablet computers, desktop computers, set-top boxes, smart TVs, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), vehicle-mounted terminals, virtual reality (VR) devices, augmented reality (AR) devices, and other types of user terminals, and can also be implemented as servers. Below, exemplary applications when the electronic device is implemented as a terminal device or a server will be described.
[0055] Referring to Figure 1 , which is a schematic diagram of an application mode of a text processing model training method provided in an embodiment of the present application, Figure 1 involves, for example, a server 200, a network 300, a terminal device 400, and a database 500. The terminal device 400 is connected to the server 200 via the network 300, which can be a wide area network or a local area network, or a combination of the two.
[0056] In some embodiments, the database 500 is used to store a large amount of text data, which can be used as data for training a text processing model. The server 200 obtains a large amount of text data from the database as training data, calls the training method of the text processing model provided in the embodiment of the present application, trains the text processing model, and obtains a trained model. The user inputs the text to be replied through the terminal device 400, and the terminal device 400 sends the text to be replied to the server 200 via the network 300. The server 200 calls the trained text processing model to perform text prediction processing on the text to be replied, obtains the reply text, and feeds the reply text back to the terminal device 400 via the network 300.
[0057] In some embodiments, the text processing model trained by the text processing model training method of the embodiments of the present application can also be used in the following application scenarios: intelligent question answering, by calling the text processing model trained by the embodiments of the present application to determine the response content of the user input question. The response content can be in various fields such as news, education, and medical care.
[0058] The embodiments of the present application can be implemented using database technology. A database, in short, can be considered an electronic filing cabinet that stores electronic files, allowing users to add, query, update, and delete data in these files. A "database" is a collection of data that is stored together in a specific manner, can be shared by multiple users, has minimal redundancy, and is independent of applications.
[0059] A database management system (DBMS) is a computer software system designed for managing databases, typically providing basic functions such as storage, retrieval, security, and backup. DBMSs can be categorized by the database model they support, such as relational or XML (Extensible Markup Language); by the type of computer they support, such as server clusters or mobile phones; by the query language they use, such as SQL or XQuery; by performance priorities, such as maximum scale or maximum speed; or by other classification methods. Regardless of the classification method used, some DBMSs are cross-category, for example, supporting multiple query languages simultaneously.
[0060] In some embodiments, the server 200 may be composed of a training server and a text processing server. The training server is used to train a text processing model, and the text processing server calls the trained model to perform text prediction.
[0061] In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal device and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of the present application.
[0062] Referring to Figure 2, Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may be the server 200 of Figure 1. The server 200 shown in Figure 2 includes: at least one processor 410, a memory 450, and at least one network interface 420. The various components in the server 200 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 440 in Figure 2.
[0063] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0064] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0065] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0066] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0067] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0068] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).
[0069] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. FIG2 shows a training device 455 for a text processing model stored in a memory 450, which can be software in the form of programs and plug-ins, including the following software modules: a data acquisition module 4551, a model training module 4552, and a text processing module 4553. These modules are logical and can be arbitrarily combined or further split according to the functions implemented. In FIG2, all the above modules are shown at once for the sake of convenience, but it should not be regarded as excluding the implementation of the text processing model training device 455 that can only include the data acquisition module 4551 and the model training module 4552. The functions of each module will be explained below.
[0070] In other embodiments, the training device of the text processing model provided in the embodiments of the present application can be implemented in hardware. As an example, the training device of the text processing model provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the text processing model provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0071] In some embodiments, the terminal or server can implement the training method of the text processing model provided in the embodiment of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a live broadcast APP or an instant messaging APP; it can also be a small program, that is, a program that can be run only by downloading it into a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be an application, module or plug-in in any form.
[0072] The training method of the text processing model provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server provided in the embodiment of the present application.
[0073] The following describes the text processing model training method provided by the embodiments of the present application. As previously mentioned, the electronic device implementing the text processing model training method of the embodiments of the present application can be a terminal or a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.
[0074] It should be noted that the examples of text processing below are explained using question-and-answer scenarios and translation scenarios as examples. Based on their understanding of the following, those skilled in the art can apply the training method of the text processing model provided in the embodiments of this application to the processing of other scenarios that require text prediction, such as: copywriting, intelligent assistants, etc.
[0075] See Figure 3A, which is a first flow chart of the training method of the text processing model provided in an embodiment of the present application, and will be explained in conjunction with the steps shown in Figure 3A.
[0076] In step 301, a sample text set is obtained.
[0077] For example, a sample text set includes: multiple independent sample texts and sample labels for multiple independent sample texts. In natural language processing (NLP), independent sample texts refer to text data that has no direct connection or dependency relationship with each other. Each text sample is an object of independent analysis and will not be affected by other text samples. The content of independent sample texts can be unrelated or related. Each independent sample text is a complete sentence or phrase. Each independent sample text represents an independent data point, and the independent sample text can be classified or analyzed through algorithms.
[0078] For example, the text in the sample text set can be text scraped from the web, and the sample labels are the output text corresponding to the individual sample texts and the label probabilities for each character in the output text. For example, in a Chinese-to-English translation scenario, the Chinese text might be "How's the weather today?" The sample labels would be the English text resulting from the translation of "How's the weather today?" and the label probabilities for each character position in the English text. Character position refers to the position of the character in the text. The output text can be pre-set or generated based on a trained model.
[0079] In step 302, at least two independent sample texts among the multiple independent sample texts are spliced together to obtain multiple spliced sample texts.
[0080] For example, the splicing process is a process of completely splicing at least two texts into one text. The object to be spliced is randomly selected, the order of the sample texts in the spliced sample text is random, and the spliced sample text includes at least two independent sample texts. The contents of the two spliced sample texts may be at least partially the same, but the two spliced sample texts are not completely the same. For example: spliced sample text 1 is composed of independent sample text 1 and independent sample text 2 spliced in order, spliced sample text 2 is composed of independent sample text 3 and independent sample text 1 spliced in order, spliced sample text 1 and spliced sample text 2 both have independent sample text 1, and the two are different as a whole.
[0081] In some embodiments, referring to FIG3B , FIG3B is a second flow chart of the training method of the text processing model provided in an embodiment of the present application; step 302 of FIG3A can be implemented through steps 3021 to 3023 of FIG3B , which are described in detail below.
[0082] In step 3021, multiple selection processes are performed on at least two independent sample texts from the multiple independent sample texts to obtain multiple selected sample combinations to be spliced.
[0083] For example, each sample combination to be spliced includes at least two independent sample texts. Each sample combination to be spliced is different, and the independent sample texts selected in each selection process are at least partially different.
[0084] For example, assuming that multiple independent sample texts are independent sample texts 1 to independent sample text N, where N is a positive integer, selection processing is performed on independent sample texts 1 to N, with the number of selections in each selection process being greater than or equal to 2, and the independent sample texts selected each time being at least partially different. For example, in the first selection process, sample combination 1 to be spliced is obtained, comprising independent sample text 1, independent sample text 3, and independent sample text 4; and in the second selection process, sample combination 2 to be spliced is obtained, comprising independent sample text 3, independent sample text 4, and independent sample text 5. In the two sample combinations to be spliced, at least some of the independent sample texts are different.
[0085] In step 3022, the following processing is performed for each sample combination to be spliced: each independent sample text in the sample combination to be spliced is randomly combined into a sample sequence.
[0086] For example, the random combination can be achieved in the following manner: the independent sample texts in the sample combination to be spliced are randomly sorted to obtain a randomly sorted sample sequence.
[0087] For example, continuing with the above example, the independent sample texts in the sample combination to be spliced are combined in random order to obtain a sample sequence [independent sample text 3, independent sample text 1, independent sample text 4].
[0088] In step 3023, each independent sample text in the sample sequence is separated by a separator to obtain a concatenated text sample.
[0089] For example, the content of independent sample texts is different and not necessarily related. Separators are used to distinguish different independent sample texts. Separators, also known as delimiters, are special characters used to separate different parts of text. Types of separators include: space, tab, return / newline, semicolon, comma, colon, period (period / full stop), question mark, quotation marks, brackets, hyphen, underline, etc.
[0090] In the embodiment of this application, the symbol <eos>For example, the separator can be <eos>, in natural language processing <eos>Represents the end of a sentence (End Of Sentence), as a label for judging termination. After adding the sample sequence of separators, the sample sequence is converted into a spliced text sample. Continuing with the example above, the spliced text sample is represented as [Independent sample text 3 <eos>Independent Sample Text 1 <eos>
[0065] Independent sample text 4. Referring to Figure 7, Figure 7 is a schematic diagram of the structure of the concatenated sample text provided in an embodiment of the present application. Samples 1 to M are randomly concatenated to obtain N samples, where M and N are positive integers, and M is greater than N. Samples are separated by separators.
[0091] In the embodiments of the present application, separators are used to distinguish independent sample texts in the concatenated text, preventing the text processing model from confusing the context, thereby focusing more on learning the features and patterns of the text. This can improve the generalization ability of the text processing model, enable it to perform better on new data, and improve the accuracy of the trained text processing model. By concatenating different samples to generate longer and more numerous samples, the computing resources required to obtain samples are saved, the content of the training samples is enriched, and the text processing model's ability to understand the context is improved, thereby improving the accuracy of the text processing model.
[0092] Continuing to refer to FIG. 3A , in step 303 , the initialized text processing model is called based on the multiple spliced sample texts to perform feature extraction processing to obtain multiple spliced sample features.
[0093] For example, the text processing model can be a transformer model (Transformer), a deep neural network model based on the self-attention mechanism. Referring to Figure 4, Figure 4 is a first structural diagram of the text processing model provided in an embodiment of the present application. The text processing model 401 includes an encoder 402 and a decoder 403, and the decoder 403 includes an attention layer 4031, a normalization layer 4032, and a linear transformation layer 4033. The encoder 402 is used to perform feature extraction processing on the text, and the feature extraction processing of the text includes the following process: converting each character in the text into a token and combining the tokens into an embedding vector (vector embeddings). The decoder 403 is used to decode the embedding vector to obtain a prediction result. The attention layer 4031 of the decoder 403 is used to call the masked attention mechanism, the normalization layer 4032 is used to normalize the output result of the attention layer 4031, and the linear transformation layer 4033 is used to convert the result of the normalization processing into a prediction result.
[0094] For example, feature extraction processing can be implemented in the following way: the text processing model performs the following processing on each spliced sample text: the encoder of the text processing model converts each character in the spliced sample text into a corresponding serial number according to the vocabulary, and obtains text features in the form of a sequence. The vocabulary stores the mapping relationship between the serial number and the character, and the text features in the form of a sequence are normalized to obtain spliced sample features in the form of a feature vector.
[0095] In step 304, the masked attention mechanism is called based on each spliced sample feature to perform prediction processing to obtain the predicted text.
[0096] For example, the masked attention mechanism involves masking the features corresponding to each independent sample text in the spliced sample features. The masked attention mechanism refers to the process of masking part of the information in the attention score matrix, which can improve the accuracy of attention calculation.
[0097] In some embodiments, referring to FIG3C , FIG3C is a third flow chart of the training method of the text processing model provided in an embodiment of the present application; step 304 of FIG3A can be implemented through steps 3041 to 3044 of FIG3C , which are described in detail below.
[0098] In step 3041 , the following processing is performed for each spliced sample feature: determining a key matrix, a value matrix, and a query matrix of the spliced sample feature.
[0099] For example, the attention mechanism is used to obtain weights and to obtain the relationship between words through a weight matrix. In the embodiment of the present application, the key, value, and query are represented in matrix form. In practical applications, the above three can also be represented by vectors. The value matrix is used to represent the content of the spliced sample features, the query matrix is used to represent the query target, and the key matrix is used to represent the queried content.
[0100] The process of attention value Attention(Q, K, V) of the above attention mechanism can be represented by the following formula (1):
[0101] Among them, softmax is a normalization function, which can make the sum of weighted probability distribution equal to 1. It is the raw score of attention, representing the similarity score, which is obtained by the dot product of the query matrix Q and the key matrix K. Is the scaling factor, which can prevent the normalized result from being too large or too small, and avoid the normalized result from being either 0 or 1. T When the attention map (N*N) is calculated.
[0102] In step 3042, a mask matrix corresponding to each independent sample text in the spliced sample features is determined.
[0103] For example, the mask matrices of different types of independent sample texts are different. That is to say, each independent sample text is masked according to its type. Referring to Figure 6, Figure 6 is a structural diagram of the attention matrix provided in an embodiment of the present application. In the attention matrix, the features of the multi-round dialogue 601 are located in the first area 603, and the features of the pan-logical thinking chain 602 are located in the second area 604. Different types of independent sample texts are masked separately. The blank part in the attention matrix is the area corresponding to the attention mask. The multi-round dialogue 601 and the pan-logical thinking chain 602 are of different types and have their own mask matrices.
[0104] For example, when the attention weight value is represented by an attention matrix, the elements outside the mask matrix in the attention matrix are combined into a sawtooth shape, and the elements located in the teeth part of the sawtooth shape are features of independent sample texts.
[0105] A matrix in which some elements form a jagged shape is a special structure, often called a "jagged matrix" or "segmented matrix." Continuing with Figure 6 , the characteristics of multi-turn conversation 601 are located in first region 603, which represents a tooth in the jagged shape. Each element in the mask matrix can be either 0 or 1.
[0106] In step 3043, the attention weight value of the spliced sample feature is determined based on the key matrix, the value matrix, the query matrix and each mask matrix.
[0107] For example, the calculation method of the key-value attention mechanism is used as an example in the embodiment of the present application. Based on the above formula (1), the mask matrix M is added to obtain the attention weight value of the spliced sample feature.
[0108] In some embodiments, referring to FIG3D , FIG3D is a fourth flow chart of the training method of the text processing model provided in an embodiment of the present application; step 3043 of FIG3A can be implemented through steps 30431 to 30434 of FIG3D , which is described in detail below.
[0109] In step 30431, the first product between the key matrix and the query matrix is obtained.
[0110] For example, based on the example of formula (1) above, the first product between the key matrix and the query matrix is represented as QK T .
[0111] In step 30432, the first product is masked based on each mask matrix to obtain a masked result.
[0112] For example, the calculation method of the key-value attention mechanism is used as an example to illustrate the first product QK T Masking is performed, and the masked result can be represented as QK T M.
[0113] In step 30433, the normalized result of the mask result is obtained.
[0114] For example, to prevent the normalized result from being too large or too small, you can set a scaling factor for the mask result. The normalized result is represented as
[0115] In step 30434, the product between the normalized result and the value matrix is used as the attention weight value of the spliced sample feature.
[0116] For example, step 30434 can be represented by the following formula (2):
[0117] Among them, Attention(Q, K, V) is the attention weight value.
[0118] Continuing to refer to FIG3C , in step 3044 , based on the attention weight value and the splicing sample features, the subsequent text of the splicing sample text is predicted to obtain the predicted text.
[0119] For example, in natural language processing, the next text usually refers to the next adjacent text segment after a given text. This concept is usually used to process coherent text sequences, such as documents, paragraphs, or continuous discourse in a conversation. In the embodiment of the present application, the next text of the spliced sample text refers to the text obtained by processing the spliced sample text according to the text processing task. Depending on the text processing task, the next text can have different contents, which are explained in detail below.
[0120] Depending on the specific application scenario of the text processing task, the subsequent text can contain different content. For example, in the English-to-Chinese translation scenario, assuming the spliced sample text is in English, the subsequent text of the spliced sample text is the Chinese equivalent of the English content. For example, in the question-and-answer scenario, assuming the spliced sample text is a question, the subsequent text of the spliced sample text is the reply. For example, in the text writing scenario, assuming the spliced sample text is a keyword or title, the subsequent text of the spliced sample text is the continuation of the keyword or title. During the prediction process, predictions are made sequentially for each character or word in the predicted text.
[0121] In some embodiments, step 3044 can be implemented by predicting multiple first prediction probabilities for each character position in the subsequent text of the spliced sample text based on the attention weight value and the spliced sample features, where each first prediction probability corresponds to a candidate character; selecting the candidate character with the highest first prediction probability corresponding to each character position as the target character; and combining each target character according to the order of each character position to obtain a predicted text. The number of predicted texts obtained during the prediction process can be at least one.
[0122] For example, the first prediction probability is the probability that the candidate character appears at the corresponding character position. The input of each i-time prediction process can be the features of the predicted target character, the attention weight value, and the spliced sample features. The output of each i-time prediction process is the first prediction probability of multiple candidate characters appearing at the i+1th character position. The candidate character corresponding to the largest first prediction probability is selected as the i+1th target character.
[0123] In an embodiment of the present application, by introducing a masked attention mechanism for text prediction processing, the accuracy of the text processing model in predicting text can be improved. The masked attention mechanism can help the text processing model focus on the local features of the text, improve the generalization ability of the text processing model, and make the text processing model more accurate on new data that has not been seen.
[0124] Continuing to refer to FIG. 3A , in step 305 , a loss function is determined based on the difference between the predicted text and the sample label.
[0125] For example, the loss function used to measure the difference between the predicted text and the sample label can be a relative entropy loss function or a vector space distance function. The vector space distance function uses the spatial distance between two sets of probability vectors as the loss value. The relative entropy loss function is used to measure the difference between two probability distributions.
[0126] In some embodiments, each sample label includes: a label probability sequence of the output text corresponding to multiple independent sample texts; step 305 can be implemented by: obtaining the first prediction probability corresponding to each character in the predicted text; combining each first prediction probability into a prediction probability sequence; representing the prediction probability sequence and the label probability sequence as vectors respectively; and using the vector space distance between the two vectors as the loss function.
[0127] For example, the label probability sequence is represented as a vector P, the prediction probability sequence is represented as a vector Q, the difference vector between vector P and vector Q is obtained, the total sum of the squares of the values of each dimension in the difference vector is obtained, the square root of the total sum is taken, and the value of the loss function is obtained.
[0128] In an embodiment of the present application, the vector space distance corresponding to the real sample label and the predicted text is used as the loss function, which improves the interpretability of the loss function, can intuitively represent the difference between the predicted data and the real data, and improves the versatility of the text processing model.
[0129] In some embodiments, each sample label includes: the label probability of each character in the output text corresponding to multiple independent sample texts; step 305 can be implemented in the following manner: obtain the first prediction probability corresponding to each character in the predicted text; perform the following processing for each character position: obtain the ratio between the label probability of the character position and the first prediction probability; obtain the second product between the logarithm of the ratio and the label probability; and use the sum of the second products of each character position as the loss function.
[0130] For example, relative entropy (KL divergence) is a (asymmetric) metric used to measure the similarity of two probability distributions. The loss function can be expressed as the following formula (3):
[0131] Among them, p(x i ) represents the distribution of label probability, q(x i ) characterizes the distribution of the first predicted probability, and the second product is D KL (p||q) is the value of the relative entropy loss function. The smaller the loss value, the smaller p(x i ) and q(x i ), the closer it is to , the better the text processing model is.
[0132] In the embodiment of the present application, relative entropy can effectively measure the difference between two probability distributions, that is, the degree of inconsistency between the true distribution and the model predicted distribution, and thus the relative entropy loss function is robust to class-imbalanced data sets. By training the text processing model through the relative entropy loss function, the generalization of the text processing model for different types of text and the accuracy of text processing can be improved.
[0133] In step 306 , the parameters of the text processing model are updated based on the loss function to obtain a trained text processing model.
[0134] For example, the parameter update process can be implemented by back propagation based on the loss function, and the parameter update process can be performed iteratively. Back propagation is an algorithm for calculating the gradient of parameters in a neural network. Back propagation uses the chain rule to calculate the gradient of the loss function for each parameter, so that gradient descent or other optimization algorithms can be used to update the network parameters, thereby minimizing the value of the loss function. Back propagation can be implemented in the following way: using the chain rule, starting from the output layer of the text processing model (which can be the output layer of the decoder of the text processing model in the embodiment of the present application), the gradient of each layer is calculated in reverse. For each weight, the gradient of the weight to the loss function is calculated in the following way: the gradient of the next input node is multiplied by the gradient from the next input node to the current layer. Repeat the above calculation process until the input layer of the text processing model is calculated. The parameters of each layer in the text processing model are updated based on the calculated gradient to obtain a trained text processing model.
[0135] In the embodiments of this application, by splicing independent sample texts to obtain spliced text as training data, the sample richness is improved and the computing resources required to obtain the sample are saved. The spliced text includes different independent sample texts, which can improve the text processing model's ability to understand context and improve the accuracy of the text processing model. During the training process, the introduction of the masked attention mechanism can improve the generalization ability of the text processing model and make the text processing model more accurate on new unseen data.
[0136] In some embodiments, after step 306, the following processing can be performed: in response to receiving the text to be replied to of the target object, obtaining the historical reply text for the target object; splicing the text to be replied to and the historical reply text into a spliced text; calling the text processing model based on the spliced text to perform text prediction processing to obtain the current reply text, wherein the current reply text is the reply content of the text to be replied to.
[0137] For example, the target object can be the terminal device or account used by the user. With the server as the execution subject, the following description is combined with a specific application scenario. Assuming that the user is not a new user, the user uses the intelligent question-and-answer platform through a terminal device. In response to receiving a question request carrying the user's corresponding account identifier and the text to be replied, the user obtains the historical question content of the account corresponding to the account identifier, splices the text to be replied and the historical reply text into a spliced text, and combines the current content to be replied and the historical reply content to feedback the reply text to the user. For example: the historical reply text is a copy with a specific language style. A new reply text is determined based on the text to be replied and the historical reply text entered by the user. The new reply text also has the above-mentioned specific language style. For another example: when the text to be replied is further question information of the historical reply text, a new reply text is determined based on the text to be replied and the historical reply text entered by the user. The new reply text is the refinement of the historical reply text based on the question information.
[0138] In an embodiment of the present application, by splicing historical reply content with the text to be replied and making full use of the above information, different segments of the reply text output by the text processing model can be correlated (for example: consistent language style, similarities in content), thereby improving the accuracy of the copy content output by the text processing model and enhancing the user experience.
[0139] In the process of training the model, the embodiment of the present application splices short independent sample texts into longer spliced texts, which improves the richness of the training samples and can save the computing resources required to obtain the training samples. The spliced text is used to improve the multi-round dialogue capability of the model and the accuracy of the training text processing model; the masked attention mechanism is used for the spliced sample features of the spliced text samples. Since there is no correlation between independent samples, each independent sample text is masked based on the masked attention mechanism, allowing the model to better generalize the language and copywriting knowledge rules, while enabling the text processing model to have better context parsing capabilities, thereby improving the accuracy of the text processing model.
[0140] Below, an exemplary application of the training method of the text processing model of the embodiment of the present application in an actual application scenario will be described.
[0141] During the supervised fine-tuning phase of existing large-scale language models, different samples are combined into the same batch for training, which leads to different attention masking mechanisms. Existing supervised fine-tuning methods for large generative language models generally use the self-attention mechanism. Due to its autoregressive nature, a token at any position in the self-attention matrix is only influenced by the token preceding it. Therefore, a lower triangular matrix masking mechanism is used to achieve this, meaning that subsequent tokens are invisible to the preceding token.
[0142] In the fine-tuning of large language models, a training batch is composed of multiple training samples, and the samples were previously independent of each other. When dealing with this situation, the existing technology generally has two solutions: one is to treat the entire batch as a complete whole, or as completely independent samples. Using a lower triangular matrix mask, when multiple samples are regarded as a single sample, the model's contextual ability will be lost, that is, the previous content will be ignored, and more attention will be paid to closer words, which will cause the multi-round dialogue ability to be weakened; the second is to form a batch attention matrix by splicing multiple small lower triangular matrix masks along the diagonal according to the number of samples. When multiple samples are regarded as completely independent samples, the attention is finely separated based on the sample granularity, which may cause the model to overfit, resulting in a decrease in the model's generalization ability and a decrease in the ability to resist context interference.
[0143] The two main existing approaches fail to optimize attention masks based on sample information, resulting in insufficient generalization or contextual understanding. Existing technologies lack suitable attention mechanisms for supervised fine-tuning of large language models. To address this issue, the key technical feature of the embodiments of this application is the use of a dynamic sawtooth attention mechanism for supervised fine-tuning of large language models.
[0144] The embodiments of the present application can effectively optimize the deficiencies of the current attention mechanism, thereby improving the effect of large-scale language model training and the accuracy of text processing. The text processing model method proposed in the embodiments of the present application has a dynamic sawtooth attention mechanism, which determines whether to separate attention according to the category of the samples in the batch. Specifically, for natural language processing language samples (such as text generation, natural language processing basic class samples), random splicing of sample sawtooth attention is performed to allow the model to better generalize the knowledge rules of language and copywriting. For logic samples (such as reasoning mathematics, multi-round dialogue samples), independent sample sawtooth attention is retained, and the model focuses on learning deep logical dependencies.
[0145] The following is an explanation of the training method of the text processing model provided in the embodiment of the present application.
[0146] For example, in order to improve training efficiency, samples will be randomly spliced between different samples in the training batch. Refer to Figure 7, which is a structural diagram of the spliced sample text provided by an embodiment of the present application. Samples 1 to M are randomly spliced to obtain N samples. Samples are separated by separators. Different samples have independent semantic spaces, that is, there should be no contextual semantic association between samples. The attention value (attention score) between independent samples is set to zero, and only the local self-attention within the independent samples is retained.
[0147] Referring to Figure 5, Figure 5 is a second structural diagram of the text processing model provided in an embodiment of the present application; Figure 5 shows a structural diagram of the module corresponding to the attention mechanism in the text processing model. Based on the input embedding vector (the concatenated text features above), a query matrix Q, a key matrix K, and a value matrix V are generated. The query matrix Q and the key matrix K are positionally encoded and multiplied. The product passes through a normalization layer 501 for normalization processing. The normalized result is multiplied with the value matrix V to obtain the output result.
[0148] The process of attention value Attention(Q, K, V) of the above attention mechanism can be represented by the following formula (1):
[0149] Among them, softmax is a normalization function, which can make the sum of weighted probability distribution equal to 1. It is the raw score of attention, representing the similarity score, which is obtained by the dot product of the query matrix Q and the key matrix K. Is the scaling factor, which can prevent the normalized result from being too large or too small, and avoid the normalized result from being either 0 or 1. T When calculating the attention map (N*N), adding different mask matrices (N*N) can implement different attention mechanisms. The following explains each attention mechanism:
[0150] (1) Using a lower triangular matrix (N*N) in the matrix of attention values, the standard attention mask can be obtained.
[0151] (2) Add the low-rank lower triangular matrix of sample granularity and splice it into a complete N*N matrix along the diagonal to obtain the sample-independent attention mask. Since the visual appearance of the low-rank lower triangular matrix after splicing is similar to a sawtooth shape, it is called the sample-independent sawtooth attention mask.
[0152] (3) According to the category information of the sample, different attention masks are used for different samples, and they are spliced into a complete N*N matrix along the diagonal to obtain a dynamic sawtooth attention mask. Refer to Figure 6, which is a structural diagram of the attention matrix provided by an embodiment of the present application. In the attention matrix, the features of the multi-round dialogue 601 are located in the first area 603, and the features of the general logical thinking chain 602 are located in the second area 604. The blank part in the attention matrix is the area corresponding to the attention mask. That is, the part outside the area corresponding to the mask forms a sawtooth shape. After adding the attention mask, the attention calculation method can be rewritten as the following formula (2):
[0153] Among them, M is the attention mask.
[0154] With the continuous development of artificial intelligence technology, intelligent assistants based on large language models have gradually become powerful assistants in life and work in the embodiments of this application. From smartphones, TVs, cars to various software applications, the application scenarios and functions of intelligent assistants are becoming increasingly rich. The improvement of the capabilities of large language models will bring about an improvement in the experience on the product side. In the large language model scenario, the text processing model trained by the training method of the text processing model proposed in the embodiment of this application may bring better generalization ability, better multi-round dialogue ability, and better logical reasoning ability. Based on the improvement of these capabilities, the intelligent assistant based on the trained text processing model can be applied in the following application scenarios:
[0155] (1) Intelligent customer service. In the field of customer service, intelligent assistants can replace traditional customer service personnel and provide users with online services 24 hours a day. Through natural language processing technology, intelligent assistants can understand users' questions and give corresponding answers. This not only improves the efficiency of customer service, but also reduces the company's labor costs. For example, intelligent customer service is used in industries such as banking, telecommunications, and e-commerce.
[0156] (2) Personal assistants. In the field of personal life, smart assistants are gradually becoming people's personal assistants. They can help users manage their schedules, remind important matters, query information, and make purchases. In addition, smart assistants can also be linked with other smart devices to achieve a more convenient personal life.
[0157] (3) Educational guidance. In the field of education, intelligent assistants can provide students with personalized learning guidance, answer questions, and improve learning efficiency. Through adaptive learning technology, intelligent assistants can recommend appropriate learning resources and exercises based on students' learning progress and abilities. In addition, intelligent assistants can also collaborate with teachers to assist them in completing student management and teaching tasks.
[0158] (4) Medical consultation. In the medical field, intelligent assistants can provide users with services such as health consultation, drug information query, and disease diagnosis. For example, by analyzing a large amount of medical literature and case studies, they can provide doctors with auxiliary diagnosis suggestions and provide patients with psychological counseling services to help relieve psychological stress.
[0159] (5) News push. In the news field, smart assistants can push personalized news information based on the user's interests and hobbies. By analyzing the user's browsing history and behavior data, smart assistants can create a unique news reading space for the user. In addition, smart assistants can also realize voice broadcasting of news, allowing users to keep up with current affairs in their busy lives.
[0160] In the embodiment of the present application, for language samples of Natural Language Processing (NLP) (such as text generation and basic natural language processing samples), random splicing of sample sawtooth attention is performed to allow the model to better generalize the knowledge rules of language and text; for logic samples (such as reasoning mathematics and multi-round dialogue samples), independent sample sawtooth attention is retained. For logic samples (such as reasoning mathematics and multi-round dialogue samples), independent sample sawtooth attention is retained, and the text processing model focuses on learning deep logical dependencies. After rigorous manual evaluation, the large-scale language text processing model trained by the method proposed in the embodiment of the present application has an overall performance in terms of basic natural language processing capabilities, multi-round dialogue capabilities, domain application capabilities, reasoning capabilities, and text generation capabilities, which are all better than the text processing model trained based on the original method.
[0161] The following continues to describe an exemplary structure of the text processing model training device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, as shown in FIG2 , the software modules in the text processing model training device 455 stored in the memory 450 may include: a data acquisition module 4551, configured to acquire a sample text set, wherein the sample text set includes: a plurality of independent sample texts and sample labels of the plurality of independent sample texts; the data acquisition module 4551, configured to splice at least two of the plurality of independent sample texts to obtain a plurality of spliced sample texts; a model training module 4552, configured to perform a splicing process on at least two of the plurality of independent sample texts to obtain a plurality of spliced sample texts; The sample text calls the initialized text processing model to perform feature extraction processing to obtain multiple spliced sample features; the model training module 4552 is configured to call the masked attention mechanism based on each of the spliced sample features to perform prediction processing to obtain predicted text, wherein the masked attention mechanism includes: masking the features corresponding to each of the independent sample texts in the spliced sample features respectively; the model training module 4552 is configured to determine the loss function based on the difference between the predicted text and the sample label; the model training module 4552 is configured to perform parameter update processing on the text processing model based on the loss function to obtain the trained text processing model.
[0162] In some embodiments, the data acquisition module 4551 is configured to perform multiple selection processes on at least two of the multiple independent sample texts to obtain multiple selected sample combinations to be spliced, wherein each of the sample combinations to be spliced includes at least two independent sample texts; perform the following processing on each of the sample combinations to be spliced: randomly combine each of the independent sample texts in the sample combination to be spliced into a sample sequence; separate each of the independent sample texts in the sample sequence with a separator to obtain a spliced text sample.
[0163] In some embodiments, the model training module 4552 is configured to perform the following processing for each of the spliced sample features: determine the key matrix, value matrix and query matrix of the spliced sample feature; determine the mask matrix corresponding to each of the independent sample texts in the spliced sample feature; determine the attention weight value of the spliced sample feature based on the key matrix, the value matrix, the query matrix and each of the mask matrices; predict the subsequent text of the spliced sample text based on the attention weight value and the spliced sample feature to obtain the predicted text.
[0164] In some embodiments, the model training module 4552 is configured to obtain a first product between the key matrix and the query matrix; mask the first product based on each mask matrix to obtain a mask result; obtain a normalized result of the mask result; and use the product between the normalized result and the value matrix as the attention weight value of the spliced sample feature.
[0165] In some embodiments, the mask matrices of different types of independent sample texts are different.
[0166] In some embodiments, when the attention weight value is represented by an attention matrix, the elements outside the mask matrix in the attention matrix are combined into a sawtooth shape, and the elements located in the teeth part of the sawtooth shape are features of independent sample text.
[0167] In some embodiments, the model training module 4552 is configured to predict multiple first prediction probabilities for each character position in the subsequent text of the spliced sample text based on the attention weight value and the spliced sample features, wherein each first prediction probability corresponds to a candidate character; the candidate character with the highest first prediction probability corresponding to each of the character positions is used as the target character; and each target character is combined according to the order of each of the character positions to obtain the predicted text.
[0168] In some embodiments, each of the sample labels includes: a label probability sequence of the output text corresponding to each of the multiple independent sample texts; a model training module 4552, configured to obtain a first prediction probability corresponding to each character in the predicted text; combining each of the first prediction probabilities into a prediction probability sequence; representing the prediction probability sequence and the label probability sequence as vectors respectively; and using the vector space distance between the two vectors as a loss function.
[0169] In some embodiments, each of the sample labels includes: the label probability of each character in the output text corresponding to the multiple independent sample texts; the model training module 4552 is configured to obtain the first prediction probability corresponding to each character in the predicted text; the following processing is performed for each character position: obtaining the ratio between the label probability of the character position and the first prediction probability; obtaining the second product between the logarithm of the ratio and the label probability; and taking the sum of the second products of each character position as the loss function.
[0170] In some embodiments, the text processing module 4553 is configured to, after performing parameter update processing on the text processing model based on the loss function to obtain the trained text processing model, obtain historical reply texts for the target object in response to receiving the text to be replied to of the target object; splice the text to be replied to and the historical reply texts into spliced text; and call the text processing model based on the spliced text to perform text prediction processing to obtain the current reply text, wherein the current reply text is the reply content of the text to be replied to.
[0171] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform the text processing model training method described in the embodiment of the present application.
[0172] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which stores computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the training method of the text processing model provided in the embodiment of the present application, for example, the training method of the text processing model shown in Figure 3A.
[0173] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0174] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0175] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0176] As an example, executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0177] To sum up, in the process of training the model through the embodiments of the present application, short independent sample texts are spliced into longer spliced texts, which improves the richness of the training samples, can save the computing resources required to obtain training samples, and use the spliced text to improve the multi-round dialogue capability of the model and the accuracy of the training text processing model; the masked attention mechanism is used for the spliced sample features of the spliced text samples. Since there is no correlation between independent samples, each independent sample text is masked based on the masked attention mechanism, so that the model can better generalize the language and copywriting knowledge rules, and at the same time enable the text processing model to have better context parsing capabilities, thereby improving the accuracy of the text processing model.
[0178] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.< / eos> < / eos> < / eos> < / eos> < / eos>
Claims
1. A method for training a text processing model, the method being performed by an electronic device, the method comprising: Acquire a sample text set, wherein the sample text set includes: a plurality of independent sample texts and sample labels of the plurality of independent sample texts; splicing at least two of the multiple independent sample texts to obtain multiple spliced sample texts; Based on the multiple splicing sample texts, the initialized text processing model is called to perform feature extraction processing to obtain multiple splicing sample features; Based on each of the spliced sample features, a masked attention mechanism is called to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking the features corresponding to each of the independent sample texts in the spliced sample features respectively; Determining a loss function based on the difference between the predicted text and the sample label; The text processing model is subjected to parameter update processing based on the loss function to obtain the trained text processing model.
2. The method according to claim 1, wherein: The step of splicing at least two of the plurality of independent sample texts to obtain a plurality of spliced sample texts includes: Performing multiple selection processes on at least two of the multiple independent sample texts to obtain multiple selected sample combinations to be spliced, wherein each of the sample combinations to be spliced includes at least two of the independent sample texts; The following processing is performed for each combination of samples to be spliced: Randomly combine each of the independent sample texts in the sample combination to be spliced into a sample sequence; Each of the independent sample texts in the sample sequence is separated by a separator to obtain a concatenated text sample.
3. The method according to claim 1, wherein: The step of calling the mask attention mechanism based on each of the spliced sample features to perform prediction processing to obtain the predicted text includes: The following processing is performed for each of the spliced sample features: Determining a key matrix, a value matrix, and a query matrix of the spliced sample features; Determine a mask matrix corresponding to each of the independent sample texts in the spliced sample features; Determining an attention weight value of the spliced sample feature based on the key matrix, the value matrix, the query matrix, and each of the mask matrices; Based on the attention weight value and the concatenated sample feature, a subsequent text of the concatenated sample text is predicted to obtain the predicted text.
4. The method according to claim 3, wherein: The determining the attention weight value of the spliced sample feature based on the key matrix, the value matrix, the query matrix and each of the mask matrices comprises: obtaining a first product between the key matrix and the query matrix; Masking the first product based on each of the mask matrices to obtain a masked result; Obtaining a normalized result of the mask result; The product of the normalized result and the value matrix is used as the attention weight value of the spliced sample feature.
5. The method according to claim 3, wherein: The mask matrices of different types of independent sample texts are different.
6. The method according to any one of claims 3 to 5, wherein: When the attention weight value is represented by an attention matrix, the elements outside the mask matrix in the attention matrix are combined into a sawtooth shape, and the elements located in the teeth part of the sawtooth shape are features of independent sample text.
7. The method according to any one of claims 1 to 6, wherein: The step of predicting a subsequent text of the spliced sample text based on the attention weight value and the spliced sample feature to obtain the predicted text includes: Based on the attention weight value and the concatenated sample feature, predict multiple first prediction probabilities of each character position in a subsequent text of the concatenated sample text, wherein each of the first prediction probabilities corresponds to a candidate character; Taking the candidate character with the highest first prediction probability corresponding to each of the character positions as the target character; Each target character is combined according to the order of each character position to obtain the predicted text.
8. The method according to any one of claims 1 to 7, wherein: Each of the sample labels includes: a label probability sequence of the output texts corresponding to the multiple independent sample texts; The determining of a loss function based on the difference between the predicted text and the sample label includes: Obtaining a first prediction probability corresponding to each character in the predicted text; combining each of the first prediction probabilities into a prediction probability sequence; Representing the prediction probability sequence and the label probability sequence as vectors respectively; The vector space distance between the two vectors is used as the loss function.
9. The method according to any one of claims 1 to 8, wherein: Each of the sample labels includes: a label probability of each character in the output text corresponding to each of the multiple independent sample texts; The determining of a loss function based on the difference between the predicted text and the sample label includes: Obtaining a first prediction probability corresponding to each character in the predicted text; For each character position, the following processing is performed: Obtaining a ratio between the label probability of the character position and the first predicted probability; Obtaining a second product between the logarithm of the ratio and the label probability; The sum of the second products of each character position is used as the loss function.
10. The method according to any one of claims 1 to 9, wherein: After performing parameter updating processing on the text processing model based on the loss function to obtain the trained text processing model, the method further includes: In response to receiving the to-be-reply text of the target object, obtaining the historical reply text for the target object; splicing the text to be replied and the historical reply text into a spliced text; The text processing model is called based on the concatenated text to perform text prediction processing to obtain a current reply text, wherein the current reply text is the reply content of the text to be replied.
11. The method according to any one of claims 1 to 10, wherein the text processing model comprises: An encoder and a decoder, wherein the encoder is used to perform the feature extraction process, and the decoder includes an attention layer, and the attention layer is used to call the mask attention mechanism for prediction processing.
12. A training device for a text processing model, the device comprising: A data acquisition module is configured to acquire a sample text set, wherein the sample text set includes: a plurality of independent sample texts and sample labels of the plurality of independent sample texts; The data acquisition module is configured to perform splicing processing on at least two of the multiple independent sample texts to obtain multiple spliced sample texts; A model training module is configured to call the initialized text processing model to perform feature extraction processing based on the multiple spliced sample texts to obtain multiple spliced sample features; The model training module is configured to call a masked attention mechanism based on each of the spliced sample features to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking the features corresponding to each of the independent sample texts in the spliced sample features respectively; The model training module is configured to determine a loss function based on the difference between the predicted text and the sample label; The model training module is configured to perform parameter update processing on the text processing model based on the loss function to obtain the trained text processing model.
13. An electronic device, comprising: A memory for storing computer executable instructions or computer programs; A processor, used to implement the text processing model training method described in any one of claims 1 to 11 when executing the computer executable instructions or computer program stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implements the training method of the text processing model described in any one of claims 1 to 11.
15. A computer program product, comprising computer executable instructions or a computer program, wherein when the computer executable instructions or the computer program is executed by a processor, the training method of the text processing model described in any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Language model training method, copywriting generation method and related equipment
CN114048289A
Text processing method, text processing device, electronic equipment and storage medium
CN115130432A
Text processing method and device, model training method and device and electronic equipment
CN116956816A
Description text generation network training method and description text generation method and device
CN116956855A
Text type determination method and device, model training method and device and electronic equipment
CN116975294A
Cited By
Information authenticity judgment method based on attention mechanism
CN120234415A
Dynamic self-adaptive semantic analysis method for electronic medical record
CN120258000A
Processing apparatus and processing method
CN120654783A
Training method of thinking type recognition model and bad thinking recognition method
CN120687918A
Multi-modal homework correction system based on image-text interlaced thinking chain
CN121388999A