Text processing model training method and device, electronic equipment and storage medium

By splicing independent sample texts and using mask attention mechanism for prediction processing, the problem of poor generalization and insufficient multi-round dialogue capabilities in the prior art is solved, and higher training sample richness and model accuracy are achieved.

CN119988960APending Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202311484312.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the training process of the supervised fine-tuning stage, existing large language models lack correlation between different training texts in a unified batch, resulting in poor generalization of text processing models and insufficient multi-round dialogue capabilities.

Method used

By obtaining the sample text collection, the independent sample text is spliced, a longer spliced ​​sample text is generated, and the mask attention mechanism is used to predict the spliced ​​sample features to optimize the loss function and parameter update process.

Benefits of technology

It improves the richness of training samples, saves the computing resources for obtaining training samples, and enhances the multi-round dialogue capabilities of the model and the accuracy of the text processing model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988960A_ABST
    Figure CN119988960A_ABST
Patent Text Reader

Abstract

The invention provides a text processing model training method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a sample text set; splicing at least two independent sample texts in the plurality of independent sample texts to obtain a plurality of spliced sample texts; calling an initialized text processing model to perform feature extraction processing based on the plurality of spliced sample texts to obtain a plurality of spliced sample features; based on each spliced sample feature, calling a mask attention mechanism to carry out prediction processing to obtain a predicted text, the mask attention mechanism comprising: respectively masking each independent sample in the spliced sample features; determining a loss function based on the difference between the predicted text and the sample tag; and performing parameter updating processing on the text processing model based on the loss function to obtain a trained text processing model. According to the method and the device, the text prediction accuracy of the text processing model obtained by training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to a training method, device, electronic device and storage medium for a text processing model. Background Art

[0002] Artificial Intelligence (AI) technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of AI generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called a large model or a basic model. After fine-tuning, it can be widely used in downstream tasks in various major directions of AI. AI software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0003] In the related art, the supervised fine-tuning stage of existing large language models usually uses models in specific fields for training. During the training process, different training texts in the same batch may not necessarily belong to relevant contexts, resulting in poor generalization or multi-round dialogue capabilities of the text processing model obtained through batch training. In the related art, there is currently no good way to improve the accuracy of the trained text processing model. Summary of the invention

[0004] The embodiments of the present application provide a text processing model training method, device, electronic device, computer-readable storage medium, and computer program product, which can improve the accuracy of text prediction by the trained text processing model.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present application provides a method for training a text processing model, the method comprising:

[0007] Acquire a sample text set, wherein the sample text set includes a plurality of independent sample texts and sample labels of the plurality of independent sample texts;

[0008] splicing at least two of the multiple independent sample texts to obtain multiple spliced ​​sample texts;

[0009] Based on the multiple splicing sample texts, the initialized text processing model is called to perform feature extraction processing to obtain multiple splicing sample features;

[0010] Based on each of the spliced ​​sample features, a masked attention mechanism is called to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking each of the independent samples in the spliced ​​sample features respectively;

[0011] Determining a loss function based on the difference between the predicted text and the sample label;

[0012] The text processing model is subjected to parameter update processing based on the loss function to obtain the trained text processing model.

[0013] The present application embodiment provides a training device for a text processing model, including:

[0014] A data acquisition module is configured to acquire a sample text set, wherein the sample text set includes a plurality of independent sample texts and sample labels of the plurality of independent sample texts;

[0015] The data acquisition module is configured to perform splicing processing on at least two of the multiple independent sample texts to obtain multiple spliced ​​sample texts;

[0016] A model training module is configured to call the initialized text processing model to perform feature extraction processing based on the multiple spliced ​​sample texts to obtain multiple spliced ​​sample features;

[0017] The model training module is configured to call a masked attention mechanism based on each of the spliced ​​sample features to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking each of the independent samples in the spliced ​​sample features respectively;

[0018] The model training module is configured to determine a loss function based on the difference between the predicted text and the sample label;

[0019] The model training module is configured to perform parameter update processing on the text processing model based on the loss function to obtain the trained text processing model.

[0020] An embodiment of the present application provides an electronic device, the electronic device comprising:

[0021] A memory for storing computer executable instructions;

[0022] The processor is used to implement the training method of the text processing model provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0023] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing the training method of the text processing model provided in the embodiment of the present application when executed by a processor.

[0024] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the training method of the text processing model provided in the embodiment of the present application is implemented.

[0025] The embodiments of the present application have the following beneficial effects:

[0026] In the process of training the model, short independent sample texts are spliced ​​into longer spliced ​​texts, which improves the richness of training samples and can save the computing resources required to obtain training samples. The use of spliced ​​texts can improve the model's multi-round dialogue capabilities and the accuracy of the training text processing model; the masked attention mechanism is used for the spliced ​​sample features of the spliced ​​text samples. Since there is no correlation between independent samples, each independent sample text is masked based on the masked attention mechanism, allowing the model to better generalize the language and copywriting knowledge rules, while enabling the text processing model to have better context parsing capabilities, thereby improving the accuracy of the text processing model. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of an application mode of the training method of the text processing model provided in an embodiment of the present application;

[0028] Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0029] FIG. 3A to FIG. 3D It is a flowchart of a training method for a text processing model provided in an embodiment of the present application;

[0030] Figure 4 It is a structural diagram of a text processing model provided in an embodiment of the present application;

[0031] Figure 5 It is an optional structural diagram of a text processing model provided in an embodiment of the present application;

[0032] Figure 6 It is a structural diagram of the attention matrix provided in an embodiment of the present application;

[0033] Figure 7 It is a structural diagram of the spliced ​​sample text provided in the embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0035] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0036] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0037] In this application, the implementation of the data capture technical solution involved (for example: obtaining text data as training data from the Internet). When the above embodiments of this application are applied to specific products or technologies, the relevant data collection (for example: receiving text to be replied uploaded by users), use and processing processes should comply with the requirements of national laws and regulations, conform to the principles of legality, legitimacy and necessity, do not involve obtaining data types prohibited or restricted by laws and regulations, and will not hinder the normal operation of the target website.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0039] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0040] 1) Transformer model is a deep neural network model based on the self-attention mechanism, which is widely used in various tasks in the field of natural language processing, such as text classification, machine translation, and question-answering systems. This model can convert input sequences into output sequences while retaining important information in the input sequence. Since the transformer model performs well in processing long texts, it has been widely used in the field of Chinese natural language processing. Compared with traditional recurrent neural networks (RNNs) and convolutional neural networks (CNNs), the Transformer model can be calculated in parallel and speed up training. It is widely used in tasks such as natural language processing, speech recognition, and image generation.

[0041] 2) Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing involves natural language, which is the language people use in daily life and is closely related to linguistic research; it also involves important technologies for model training in the fields of computer science, mathematics, and artificial intelligence. The pre-trained model is developed from the Large Language Model (LLM) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0042] 3) Large Language Model (LLM) refers to a deep learning model trained with a large amount of text data that can generate natural language text or understand the meaning of language text. Large language models can handle a variety of natural language tasks, such as text classification, question answering, and dialogue, and are an important path to artificial intelligence.

[0043] 4) Normalization (Softmax) function, which is used to convert the output values ​​of different categories into a function with a probability distribution ranging from [0, 1] and 1. The formula is as follows: Among them, Z i is the output value of the ith node, and C is the number of output nodes, that is, the number of classification categories.

[0044] 5) Attention mechanism. The attention mechanism in deep learning is essentially similar to the selective visual attention mechanism of humans. The core goal of the attention mechanism is to select the information that is more critical to the current task goal from a large amount of information.

[0045] The embodiments of the present application provide a text processing model training method, a text processing model training device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of text prediction by the trained text processing model.

[0046] The following describes an exemplary application of an electronic device provided by an embodiment of the present application. The electronic device provided by an embodiment of the present application can implement a terminal device, such as a laptop, a tablet computer, a desktop computer, a set-top box, a smart TV, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), a vehicle terminal, a virtual reality (VR) device, an augmented reality (AR) device, and other types of user terminals, and can also be implemented as a server. Below, an exemplary application when the electronic device is implemented as a terminal device or a server will be described.

[0047] refer to Figure 1 , Figure 1 Schematic diagram of the application mode of the training method of the text processing model provided in the embodiment of the present application; for example, Figure 1 The server 200, network 300, terminal device 400 and database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.

[0048] In some embodiments, the database 500 is used to store a large amount of text data, and the text data can be used as data for training a text processing model. The server 200 obtains a large amount of text data from the database as training data, calls the training method of the text processing model provided in the embodiment of the present application, trains the text processing model, and obtains a trained model. The user inputs the text to be replied through the terminal device 400, and the terminal device 400 sends the text to be replied to the server 200 through the network 300. The server 200 calls the trained text processing model to perform text prediction processing on the text to be replied, obtains the reply text, and feeds the reply text back to the terminal device 400 through the network 300.

[0049] In some embodiments, the text processing model trained by the text processing model training method of the embodiment of the present application can also be applied in the following application scenarios: intelligent question answering, by calling the text processing model trained by the embodiment of the present application to determine the reply content of the question input by the user. The reply content can be in various fields such as news, education, and medical treatment.

[0050] The embodiments of the present application can be implemented through blockchain technology. The text processing model trained by the embodiments of the present application can be uploaded to the blockchain for storage, and the reliability of the text processing model can be guaranteed by the consensus algorithm. Blockchain is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains information about a batch of model data, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0051] The embodiments of the present application can be implemented through database technology. In short, a database can be regarded as an electronic file cabinet where electronic files are stored. Users can add, query, update, delete, etc. data in the files. The so-called "database" is a collection of data that is stored together in a certain way, can be shared with multiple users, has as little redundancy as possible, and is independent of the application program.

[0052] A database management system (DBMS) is a computer software system designed for managing databases. It generally has basic functions such as storage, retrieval, security, and backup. Database management systems can be classified according to the database model they support, such as relational, XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters, mobile phones; or according to the query language used, such as Structured Query Language (SQL), XQuery; or according to performance focus, such as maximum scale, maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMS can cross categories, for example, supporting multiple query languages ​​at the same time.

[0053] The embodiments of the present application can also be implemented through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application. It can form a resource pool, which is used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, as well as the promotion of search services, social networks, mobile commerce and open collaboration, each item may have its own hash code identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0054] In some embodiments, the server 200 may be composed of a training server and a text processing server. The training server is used to train a text processing model, and the text processing server calls the trained model to perform text prediction.

[0055] In some embodiments, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0056] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may be Figure 1 Server 200, Figure 2 The server 200 shown includes: at least one processor 410, a memory 450, and at least one network interface 420. The various components in the server 200 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .

[0057] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0058] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0059] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0060] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0061] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0062] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;

[0063] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 The text processing model training device 455 stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, including the following software modules: data acquisition module 4551, model training module 4552, text processing module 4553. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. Figure 2For the sake of convenience, all the above modules are shown at once, but it should not be considered that the training device 455 of the text processing model excludes the implementation that can only include the data acquisition module 4551 and the model training module 4552. The functions of each module will be explained below.

[0064] In other embodiments, the training device of the text processing model provided in the embodiments of the present application can be implemented in hardware. As an example, the training device of the text processing model provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the text processing model provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.

[0065] In some embodiments, the terminal or server can implement the training method of the text processing model provided in the embodiment of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a live broadcast APP or an instant messaging APP; it can also be a small program, that is, a program that can be run only by downloading it to a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be an application, module or plug-in in any form.

[0066] The training method of the text processing model provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server provided in the embodiment of the present application.

[0067] The following describes the training method of the text processing model provided in the embodiment of the present application. As mentioned above, the electronic device for implementing the training method of the text processing model in the embodiment of the present application can be a terminal or a server, or a combination of the two. Therefore, the execution subject of each step will not be repeatedly described below.

[0068] It should be noted that the examples of text processing below are explained using question-and-answer scenarios and translation scenarios as examples. Based on their understanding of the following text, those skilled in the art can apply the training method of the text processing model provided in the embodiments of the present application to the processing of other scenarios that require text prediction, such as: copywriting, intelligent assistants, etc.

[0069] See also Figure 3A , Figure 3A is a flowchart of a training method for a text processing model provided in an embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0070] In step 301, a sample text set is obtained.

[0071] For example, the sample text set includes multiple independent sample texts and sample labels of multiple independent sample texts. The independent sample texts may be unrelated or related, and each independent sample text is a complete sentence or phrase.

[0072] For example, the text in the sample text set can be text captured from the Internet, and the sample label is the label probability corresponding to each character in the predicted text corresponding to the text. For example: the application scenario of text processing is the translation scenario of Chinese to English. The Chinese is "How is the weather today?" The sample label is based on the English text translated from "How is the weather today?", and the label probability of each character position in the English text is the label probability of the character corresponding to each character position in the English text. The character position refers to the position of the character in the text.

[0073] In step 302, at least two independent sample texts among the multiple independent sample texts are spliced ​​to obtain multiple spliced ​​sample texts.

[0074] For example, the splicing process is performed randomly, the order of the sample texts in the spliced ​​sample text is random, and the spliced ​​sample text includes at least two independent sample texts. The contents of the two spliced ​​sample texts may be at least partially the same, but the two spliced ​​sample texts are not completely the same. For example: Spliced ​​sample text 1 is composed of independent sample text 1 and independent sample text 2 spliced ​​in sequence, and spliced ​​sample text 2 is composed of independent sample text 3 and independent sample text 1 spliced ​​in sequence. Spliced ​​sample text 1 and spliced ​​sample text 2 both have independent sample text 1, and the two are different as a whole.

[0075] In some embodiments, reference Figure 3B , Figure 3B It is a flowchart of a training method for a text processing model provided in an embodiment of the present application; Figure 3A Step 302 can be accomplished by Figure 3B Steps 3021 to 3023 are implemented as described in detail below.

[0076] In step 3021, multiple selection processes are performed on at least two independent sample texts among the multiple independent sample texts to obtain multiple selected sample combinations to be spliced.

[0077] For example, each sample combination to be spliced ​​includes at least two independent sample texts. Each sample combination to be spliced ​​is different, and the independent sample texts selected in each selection process are at least partially different.

[0078] For example, assuming that the multiple independent sample texts are independent sample text 1 to independent sample text N, N is a positive integer, and selection processing is performed on independent sample text 1 to independent sample text N, and the number of selections for each selection processing is greater than or equal to 2, and the independent sample texts selected each time are at least partially different. For example: the first selection processing obtains sample combination 1 to be spliced, including independent sample text 1, independent sample text 3, and independent sample text 4; the second selection processing obtains sample combination 2 to be spliced, including independent sample text 3, independent sample text 4, and independent sample text 5. In the two sample combinations to be spliced, at least some of the independent sample texts are different.

[0079] In step 3022, the following processing is performed for each sample combination to be spliced: each independent sample text in the sample combination to be spliced ​​is randomly combined into a sample sequence.

[0080] For example, the random combination can be achieved in the following way: randomly sorting the independent sample texts in the sample combination to be spliced ​​to obtain a randomly sorted sample sequence.

[0081] For example, continuing with the above example, the independent sample texts in the sample combination to be spliced ​​are combined in random order to obtain a sample sequence [independent sample text 3, independent sample text 1, independent sample text 4].

[0082] In step 3023, each independent sample text in the sample sequence is separated by a separator to obtain a concatenated text sample.

[0083] For example, the contents of independent sample texts are different, and the contents are not necessarily related. Different independent sample texts are distinguished by separators. The separator is also called a delimiter.

[0084] For example: The content of the separator can be <eos> , <eos>Represents the end of a sequence and serves as a label for determining termination. The sample sequence with the separator added is the concatenated text sample. Continuing with the example above, the concatenated text sample is characterized as [independent sample text 3 <eos>Independent sample text 1 <eos>Independent sample text 4]. Figure 7 , Figure 7 1 is a schematic diagram of the structure of the spliced ​​sample text provided in the embodiment of the present application. Samples 1 to M are randomly spliced ​​to obtain N samples, where M and N are positive integers, and M is greater than N. The samples are separated by separators.

[0085] In the embodiment of the present application, the independent sample texts in the spliced ​​text are distinguished by separators to avoid the text processing model from confusing the context, which can improve the accuracy of training. By splicing different samples to generate longer and more samples, the computing resources required to obtain samples are saved, the content of the training samples is enriched, and the ability of the text processing model to understand the context can be improved, thereby improving the accuracy of the text processing model.

[0086] Continue to refer Figure 3A In step 303, the initialized text processing model is called based on the multiple splicing sample texts to perform feature extraction processing to obtain multiple splicing sample features.

[0087] For example, the text processing model can be a transformer model, a deep neural network model based on the self-attention mechanism. Figure 4 , Figure 4 4 is a structural diagram of a text processing model provided in an embodiment of the present application. The text processing model 401 includes an encoder 402 and a decoder 403, and the decoder 403 includes an attention layer 4031, a normalization layer 4032, and a linear transformation layer 4033. The encoder 402 is used to perform feature extraction processing on the text, and the feature extraction processing of the text includes the following process: converting each character in the text into a token and combining the tokens into an embedding vector. The decoder 403 is used to decode the embedding vector to obtain a prediction result.

[0088] For example, feature extraction processing can be implemented in the following ways: the text processing model performs the following processing on each spliced ​​sample text: the encoder of the text processing model converts each character in the spliced ​​sample text into a corresponding serial number according to a vocabulary, and obtains text features in a sequence form. The vocabulary stores the mapping relationship between serial numbers and characters. The text features in the sequence form are normalized to obtain spliced ​​sample features in the form of feature vectors.

[0089] In step 304, the masked attention mechanism is called to perform prediction processing based on each concatenated sample feature to obtain the predicted text.

[0090] For example, the masked attention mechanism includes: masking each independent sample in the concatenated sample features. The masked attention mechanism refers to the process of masking part of the information in the attention score matrix to improve the accuracy of attention calculation.

[0091] In some embodiments, reference Figure 3C , Figure 3C It is a flowchart of a training method for a text processing model provided in an embodiment of the present application; Figure 3A Step 304 can be accomplished by Figure 3C Steps 3041 to 3044 are implemented as described in detail below.

[0092] In step 3041 , the following processing is performed for each concatenated sample feature: determining a key matrix, a value matrix, and a query matrix of the concatenated sample feature.

[0093] For example, the attention mechanism is used to obtain weights, and is used to obtain the relationship between words through a weight matrix. In the embodiment of the present application, the key, value, and query are represented in matrix form. In practical applications, the above three can also be represented by vectors. The value matrix is ​​used to represent the content of the spliced ​​sample features, the query matrix is ​​used to represent the query target, and the key matrix is ​​used to represent the queried content.

[0094] The process of attention value Attention(Q, K, V) of the above attention mechanism can be represented by the following formula (1):

[0095]

[0096] Among them, softmax is a normalization function, and softmax can make the sum of weighted probability distribution equal to 1. It is the raw score of attention, representing the similarity score, which is obtained by the dot product of the query matrix Q and the key matrix K. is the scaling factor, which can prevent the normalized result from being too large or too small, and avoid the normalized result being either 0 or 1. T When the attention map (N*N) is calculated.

[0097] In step 3042, a mask matrix corresponding to each independent sample text in the concatenated sample feature is determined.

[0098] For example, the mask matrices of different types of independent sample texts are different. That is, each independent sample text is masked according to its type. Figure 6 , Figure 6 It is a structural diagram of the attention matrix provided by the embodiment of the present application. In the attention matrix, the characteristic representation of the multi-round dialogue 601 is the first area 603, the characteristic representation of the general logical thinking chain 602 is 604, and different types of independent sample texts are masked separately. The blank part in the attention matrix is ​​the area corresponding to the attention mask. The multi-round dialogue 601 and the general logical thinking chain 602 are of different types, and each has its own mask matrix.

[0099] For example, when the attention weight value is represented by the attention matrix, the part outside the mask matrix in the attention matrix is ​​a sawtooth shape, and each tooth in the sawtooth corresponds to the feature of an independent sample text. Figure 6 The characteristic of the multi-round dialogue 601 is represented by the first region 603, and the first region 603 is a tooth in the sawtooth shape. Each element in the mask matrix can be 0 or 1.

[0100] In step 3043, the attention weight values ​​of the concatenated sample features are determined based on the key matrix, the value matrix, the query matrix and each mask matrix.

[0101] For example, the calculation method of the key-value attention mechanism is used as an example in the embodiment of the present application. Based on the above formula (1), the mask matrix M is added to obtain the attention weight value of the spliced ​​sample feature.

[0102] In some embodiments, reference Figure 3D , Figure 3D It is a flowchart of a training method for a text processing model provided in an embodiment of the present application; Figure 3A Step 3043 can be Figure 3D Steps 30431 to 30434 are implemented as described in detail below.

[0103] In step 30431, the first product between the key matrix and the query matrix is ​​obtained.

[0104] For example, based on the example of formula (1) above, the first product between the key matrix and the query matrix is ​​represented as QK T .

[0105] In step 30432, the first product is masked based on each mask matrix to obtain a masked result.

[0106] For example, the calculation method of the key-value attention mechanism is used as an example in the embodiment of the present application to illustrate the first product QK T Masking is performed, and the masked result can be represented as QK T M.

[0107] In step 30433, the normalized result of the mask result is obtained.

[0108] For example, to prevent the normalized result from being too large or too small, you can set a scaling factor for the mask result. The normalized result is represented as

[0109] In step 30434, the product of the normalized result and the value matrix is ​​used as the attention weight value of the spliced ​​sample feature.

[0110] For example, step 30434 can be represented by the following formula (2):

[0111]

[0112] Among them, Attention(Q, K, V) is the attention weight value.

[0113] Continue to refer Figure 3C In step 3044, based on the attention weight value and the concatenated sample features, the next sentence of the concatenated sample text is predicted to obtain the predicted text.

[0114] For example, depending on the specific application scenario, the next sentence of the text can be different. Taking the English-to-Chinese translation scenario as an example, assuming that the spliced ​​sample text is English, the next sentence of the spliced ​​sample text is the Chinese content corresponding to the English content; taking the question-and-answer scenario as an example, assuming that the spliced ​​sample text is a question, the next sentence of the spliced ​​sample text is the reply content; taking the text writing scenario as an example, the spliced ​​sample text is a keyword or a title, and the next sentence of the spliced ​​sample text is the continuation of the keyword or title. During the prediction process, predictions are made for each character or each word in the predicted text.

[0115] In some embodiments, step 3044 can be implemented in the following manner: based on the attention weight value and the concatenated sample features, predict multiple first prediction probabilities for each character position in the next sentence text of the concatenated sample text, wherein each first prediction probability corresponds to a candidate character; the candidate character with the highest first prediction probability corresponding to each character position is used as the target character; each target character is combined according to the order of each character position to obtain a predicted text. The number of predicted texts obtained in the prediction process can be at least one.

[0116] Continue to refer Figure 3A , in step 305, a loss function is determined based on the difference between the predicted text and the sample label.

[0117] For example, the loss function that measures the difference between the predicted text and the sample label can be a relative entropy loss function or a vector space distance function. The vector space distance function uses the spatial distance between two sets of probability vectors as the loss value. The relative entropy loss function is used to measure the difference between two probability distributions.

[0118] In some embodiments, each sample label includes: a label probability sequence of the output text corresponding to multiple independent sample texts; step 305 can be implemented in the following manner: obtain the first prediction probability corresponding to each character in the predicted text; combine each first prediction probability into a prediction probability sequence; represent the prediction probability sequence and the label probability sequence as vectors respectively; and use the vector space distance between the two vectors as the loss function.

[0119] For example, the label probability sequence is represented as a vector P, the prediction probability sequence is represented as a vector Q, the difference vector between the vector P and the vector Q is obtained, the total sum of the squares of the values ​​of each dimension in the difference vector is obtained, the square root of the total sum is taken, and the value of the loss function is obtained.

[0120] In some embodiments, each sample label includes: the label probability of each character in the output text corresponding to multiple independent sample texts; step 305 can be implemented in the following manner: obtain the first prediction probability corresponding to each character in the predicted text; perform the following processing for each character position: obtain the ratio between the label probability of the character position and the first prediction probability; obtain the second product between the logarithm of the ratio and the label probability; and use the sum of the second products of each character position as the loss function.

[0121] For example, relative entropy (KL divergence) is a (asymmetric) metric used to measure the similarity of two probability distributions. The loss function can be expressed as the following formula (3):

[0122]

[0123] Among them, p(x i ) represents the distribution of label probability, q(x i ) characterizes the distribution of the first prediction probability, and the second product is D KL (p||q) is the value of the relative entropy loss function. The smaller the loss value, the smaller p(x i ) and q(x i ), the closer is the value of , the better is the effect of the text processing model.

[0124] In step 306, the parameters of the text processing model are updated based on the loss function to obtain a trained text processing model.

[0125] For example, the parameter update process can be implemented by back propagation based on the loss function, and the parameter update process can be performed iteratively.

[0126] In some embodiments, after step 306, the following processing can be performed: in response to receiving the text to be replied to of the target object, obtaining the historical reply text for the target object; splicing the text to be replied to and the historical reply text into a spliced ​​text; calling a text processing model based on the spliced ​​text to perform text prediction processing to obtain the current reply text, wherein the current reply text is the reply content of the text to be replied to.

[0127] For example, the target object may be a terminal device or account used by the user. The server is used as the execution subject, and the following is explained in combination with a specific application scenario. Assuming that the user is not a new user, the user uses the intelligent question-and-answer platform through a terminal device, and in response to receiving a question request carrying the user's corresponding account identifier and the text to be replied, obtains the historical question content of the account corresponding to the account identifier, splices the text to be replied and the historical reply text into a spliced ​​text, and feeds back the reply text to the user in combination with the current content to be replied and the historical reply content. For example: the historical reply text is a copy with a specific language style, and a new reply text is determined based on the text to be replied and the historical reply text entered by the user, and the new reply text also has the above-mentioned specific language style. For another example: when the text to be replied is further question information of the historical reply text, a new reply text is determined based on the text to be replied and the historical reply text entered by the user, and the new reply text is a refinement of the historical reply text based on the question information.

[0128] In an embodiment of the present application, by splicing the historical reply content with the text to be replied to and making full use of the above information, different segments of the reply text output by the text processing model can be related (for example: consistent language style, similarities in content), thereby improving the accuracy of the copy content output by the text processing model and improving the user experience.

[0129] In the process of training the model, the embodiment of the present application splices short independent sample texts into longer spliced ​​texts, thereby improving the richness of training samples and saving computing resources required for obtaining training samples. The spliced ​​text is used to improve the multi-round dialogue capability of the model and the accuracy of the training text processing model. A masked attention mechanism is used for the spliced ​​sample features of the spliced ​​text samples. Since there is no correlation between independent samples, each independent sample text is masked based on the masked attention mechanism, so that the model can better generalize the language and copywriting knowledge rules, and at the same time, the text processing model has better context parsing capabilities, thereby improving the accuracy of the text processing model.

[0130] Below, an exemplary application of the training method of the text processing model of the embodiment of the present application in an actual application scenario will be described.

[0131] In the supervised fine-tuning stage of existing large language models, different samples are spliced ​​into the same batch for training, which brings about different attention mask mechanisms. In the existing supervised fine-tuning methods of large generative language models, the self-attention mechanism is generally used. Due to its autoregressive generation characteristics, the token at any position in the self-attention matrix is ​​only affected by the token before it. Therefore, the lower triangular matrix mask mechanism is used to achieve this feature, that is, the subsequent token is invisible to the previous token.

[0132] In the fine-tuning of large language models, a training batch is composed of multiple training samples, and the samples were previously independent of each other. There are generally two solutions for dealing with this situation in the prior art: one is to treat the entire batch as a complete whole, or as completely independent samples. Using a lower triangular matrix mask, when multiple samples are regarded as a single sample, the model's contextual ability will be lost, that is, the previous content will be ignored, and more attention will be paid to closer words, which will cause the weakening of multi-round dialogue capabilities; the second is to form a batch of attention matrices by splicing multiple small lower triangular matrix masks along the diagonal according to the number of samples. When multiple samples are regarded as completely independent samples, the attention is finely separated based on the sample granularity, which may cause the model to overfit, resulting in a decrease in the model's generalization ability and a decrease in the ability to resist context interference.

[0133] The two main existing methods cannot perform targeted attention mask optimization based on sample information, resulting in insufficient generalization or context understanding of the model. The prior art lacks a suitable attention mechanism in the supervised fine-tuning of large language models. To address this problem, the key technical point of the embodiment of the present application uses a dynamic sawtooth attention mechanism to supervise fine-tune large language models.

[0134] The embodiments of the present application can effectively optimize the shortcomings of the current attention mechanism, thereby improving the effect of large-scale language model training and the accuracy of text processing. The text processing model method proposed in the embodiments of the present application has a dynamic sawtooth attention mechanism, which determines whether to separate attention according to the category of the samples in the batch. Specifically, for natural language processing language samples (such as text generation, natural language processing basic class samples), the sample sawtooth attention is randomly spliced ​​to allow the model to better generalize the knowledge rules of language and copywriting. For logic samples (such as reasoning mathematics, multi-round dialogue samples), the independent sample sawtooth attention is retained, and the model focuses on learning deep logical dependencies.

[0135] The following is an explanation of the training method of the text processing model provided in the embodiment of the present application.

[0136] For example, in order to improve training efficiency, samples in different training batches will be randomly spliced. Figure 7 , Figure 7 : It is a structural diagram of the spliced ​​sample text provided in the embodiment of the present application. Samples 1 to M are randomly spliced ​​to obtain N samples. Samples are separated by separators. Different samples have independent semantic spaces, that is, there should be no contextual semantic association between samples. The attention score between independent samples is set to zero, and only the local self-attention within the independent samples is retained.

[0137] refer to Figure 5 , Figure 5 It is an optional structural diagram of a text processing model provided in an embodiment of the present application; Figure 5 The structural diagram of the module corresponding to the attention mechanism in the text processing model is shown. Based on the input embedding vector (the concatenated text feature above), the query matrix Q, the key matrix K, and the value matrix V are generated. The query matrix Q and the key matrix K are respectively positionally encoded and multiplied. The product passes through the normalization layer 501 for normalization processing. The normalized result is multiplied with the value matrix V to obtain the output result.

[0138] The process of attention value Attention(Q, K, V) of the above attention mechanism can be represented by the following formula (1):

[0139]

[0140] Among them, softmax is a normalization function, and softmax can make the sum of weighted probability distribution equal to 1. It is the raw score of attention, representing the similarity score, which is obtained by the dot product of the query matrix Q and the key matrix K. is the scaling factor, which can prevent the normalized result from being too large or too small, and avoid the normalized result being either 0 or 1. T When the attention map (N*N) is calculated, different attention mechanisms can be implemented by adding different mask matrices (N*N). The following explains each attention mechanism:

[0141] (1) Using a lower triangular matrix (N*N) in the matrix of attention values, we can get the standard attention mask.

[0142] (2) Add a low-rank lower triangular matrix of sample granularity and splice it along the diagonal to form a complete N*N matrix to obtain a sample-independent attention mask. Since the visual appearance of the low-rank lower triangular matrix after splicing is similar to a jagged shape, it is called a sample-independent jagged attention mask.

[0143] (3) According to the category information of the sample, different attention masks are used for different samples, and they are spliced ​​along the diagonal to form a complete N*N matrix to obtain a dynamic sawtooth attention mask. Figure 6 , Figure 6 601 is a schematic diagram of the structure of the attention matrix provided by the embodiment of the present application. In the attention matrix, the characteristic representation of the multi-round dialogue 601 is the first area 603, and the characteristic representation of the general logic thinking chain 602 is 604. The blank part in the attention matrix is ​​the area corresponding to the attention mask. That is, the part outside the area corresponding to the mask forms a zigzag shape. After adding the attention mask, the attention calculation method can be rewritten as the following formula (2):

[0144]

[0145] Among them, M is the attention mask.

[0146] With the continuous development of artificial intelligence technology, intelligent assistants based on large language models have gradually become powerful assistants in life and work in the embodiments of this application. From smartphones, televisions, cars to various software applications, the application scenarios and functions of intelligent assistants are becoming increasingly rich. The improvement of the capabilities of large language models will bring about an improvement in the experience on the product side. In the large language model scenario, the text processing model trained by the training method of the text processing model proposed in the embodiment of this application may bring better generalization capabilities, better multi-round dialogue capabilities, and better logical reasoning capabilities. Based on the improvement of these capabilities, intelligent assistants based on trained text processing models can be applied in the following application scenarios:

[0147] (1) Intelligent customer service. In the field of customer service, intelligent assistants can replace traditional customer service personnel and provide users with 24-hour online services. Through natural language processing technology, intelligent assistants can understand users' questions and give corresponding answers. This not only improves the efficiency of customer service, but also reduces the labor costs of enterprises. For example, intelligent customer service is used in banking, telecommunications, e-commerce and other industries.

[0148] (2) Personal assistants. In the field of personal life, smart assistants are gradually becoming people's personal assistants. They can help users manage their schedules, remind important matters, query information, and make purchases. In addition, smart assistants can also be linked with other smart devices to achieve a more convenient personal life.

[0149] (3) Educational guidance. In the field of education, intelligent assistants can provide students with personalized learning guidance, answer questions and improve learning efficiency. Through adaptive learning technology, intelligent assistants can recommend appropriate learning resources and exercises based on students' learning progress and abilities. In addition, intelligent assistants can also collaborate with teachers to assist them in completing student management and teaching tasks.

[0150] (4) Medical consultation. In the medical field, intelligent assistants can provide users with services such as health consultation, drug information query, and disease diagnosis. For example, by analyzing a large amount of medical literature and cases, they can provide doctors with auxiliary diagnosis suggestions and provide patients with psychological counseling services to help relieve psychological stress.

[0151] (5) News push. In the field of news, smart assistants can push personalized news information based on the user's interests and hobbies. By analyzing the user's browsing history and behavior data, smart assistants can create a unique news reading space for the user. In addition, smart assistants can also realize voice broadcasting of news, allowing users to keep up with current affairs in their busy lives.

[0152] In the embodiments of the present application, for language samples of natural language processing (NLP) (such as text generation and basic natural language processing samples), random splicing of sample sawtooth attention is performed to allow the model to better generalize the knowledge rules of language and copywriting; for logic samples (such as reasoning mathematics and multi-round dialogue samples), independent sample sawtooth attention is retained. For logic samples (such as reasoning mathematics and multi-round dialogue samples), independent sample sawtooth attention is retained, and the text processing model focuses on learning deep logical dependencies. After rigorous manual evaluation, the large-scale language text processing model trained by the method proposed in the embodiments of the present application has an overall performance in terms of basic natural language processing capabilities, multi-round dialogue capabilities, domain application capabilities, reasoning capabilities, and text generation capabilities, which are all better than the text processing model trained based on the original method.

[0153] The following is a description of an exemplary structure of a text processing model training device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules in the training device 455 of the text processing model stored in the memory 450 may include: a data acquisition module 4551, configured to acquire a sample text set, wherein the sample text set includes multiple independent sample texts and sample labels of the multiple independent sample texts; the data acquisition module 4551 is configured to perform splicing processing on at least two of the multiple independent sample texts to obtain multiple spliced ​​sample texts; a model training module 4552 is configured to call the initialized text processing model based on the multiple spliced ​​sample texts to perform feature extraction processing to obtain multiple spliced ​​sample features; the model training module 4552 is configured to call the masked attention mechanism based on each of the spliced ​​sample features to perform prediction processing to obtain predicted text, wherein the masked attention mechanism includes: masking each of the independent samples in the spliced ​​sample features respectively; the model training module 4552 is configured to determine the loss function based on the difference between the predicted text and the sample label; the model training module 4552 is configured to perform parameter update processing on the text processing model based on the loss function to obtain the trained text processing model.

[0154] In some embodiments, the data acquisition module 4551 is configured to perform multiple selection processes on at least two of the multiple independent sample texts to obtain multiple selected sample combinations to be spliced, wherein each of the sample combinations to be spliced ​​includes at least two independent sample texts; the following processes are performed for each of the sample combinations to be spliced: each of the independent sample texts in the sample combination to be spliced ​​is randomly combined into a sample sequence; each of the independent sample texts in the sample sequence is separated by a separator to obtain a spliced ​​text sample.

[0155] In some embodiments, the model training module 4552 is configured to perform the following processing for each of the spliced ​​sample features: determine the key matrix, value matrix and query matrix of the spliced ​​sample feature; determine the mask matrix corresponding to each of the independent sample texts in the spliced ​​sample feature; determine the attention weight value of the spliced ​​sample feature based on the key matrix, the value matrix, the query matrix and each of the mask matrices; and predict the next sentence of the spliced ​​sample text based on the attention weight value and the spliced ​​sample feature to obtain the predicted text.

[0156] In some embodiments, the model training module 4552 is configured to obtain a first product between the key matrix and the query matrix; mask the first product based on each of the mask matrices to obtain a masked result; obtain a normalized result of the masked result; and use the product of the normalized result and the value matrix as the attention weight value of the spliced ​​sample feature.

[0157] In some embodiments, the mask matrices of different types of the independent sample texts are different.

[0158] In some embodiments, when the attention weight value is represented by an attention matrix, the portion outside the mask matrix in the attention matrix is ​​a sawtooth shape, and each tooth in the sawtooth shape corresponds to a feature of an independent sample text.

[0159] In some embodiments, the model training module 4552 is configured to predict multiple first prediction probabilities for each character position in the next sentence text of the spliced ​​sample text based on the attention weight value and the spliced ​​sample features, wherein each of the first prediction probabilities corresponds to a candidate character; the candidate character with the highest first prediction probability corresponding to each of the character positions is taken as the target character; and each target character is combined according to the order of each of the character positions to obtain the predicted text.

[0160] In some embodiments, each of the sample labels includes: a label probability sequence of the output texts corresponding to the multiple independent sample texts; a model training module 4552 configured to obtain a first prediction probability corresponding to each character in the predicted text; combine each of the first prediction probabilities into a prediction probability sequence; represent the prediction probability sequence and the label probability sequence as vectors respectively; and use the vector space distance between the two vectors as a loss function.

[0161] In some embodiments, each of the sample labels includes: the label probability of each character in the output text corresponding to the multiple independent sample texts; a model training module 4552, configured to obtain a first prediction probability corresponding to each character in the predicted text; performing the following processing for each character position: obtaining the ratio between the label probability of the character position and the first prediction probability; obtaining the second product between the logarithm of the ratio and the label probability; and taking the sum of the second products of each character position as a loss function.

[0162] In some embodiments, the text processing module 4553 is configured to, after performing parameter update processing on the text processing model based on the loss function to obtain the trained text processing model, in response to receiving the text to be replied of the target object, obtain the historical reply text for the target object; splice the text to be replied and the historical reply text into a spliced ​​text; and call the text processing model based on the spliced ​​text to perform text prediction processing to obtain the current reply text, wherein the current reply text is the reply content of the text to be replied.

[0163] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or the computer executable instruction from the computer-readable storage medium, and the processor executes the computer program or the computer executable instruction, so that the electronic device executes the training method of the text processing model described in the embodiment of the present application.

[0164] The present application embodiment provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the text processing model training method provided by the present application embodiment, for example, Figure 3A The training method of the text processing model is shown.

[0165] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0166] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0167] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0168] As an example, the executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.

[0169] To summarize, in the process of training the model through the embodiments of the present application, short independent sample texts are spliced ​​into longer spliced ​​texts, which improves the richness of training samples and can save computing resources required for obtaining training samples. The spliced ​​text is used to improve the multi-round dialogue capability of the model and the accuracy of the training text processing model; a masked attention mechanism is used for the spliced ​​sample features of the spliced ​​text samples. Since there is no correlation between independent samples, each independent sample text is masked based on the masked attention mechanism, so that the model can better generalize the language and copywriting knowledge rules, and at the same time, the text processing model has better context parsing capabilities, thereby improving the accuracy of the text processing model.

[0170] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.< / eos> < / eos> < / eos> < / eos>

Claims

1. A training method for a text processing model, characterized in that: The method comprises: Acquire a sample text set, wherein the sample text set includes a plurality of independent sample texts and sample labels of the plurality of independent sample texts; splicing at least two of the multiple independent sample texts to obtain multiple spliced ​​sample texts; Based on the multiple splicing sample texts, the initialized text processing model is called to perform feature extraction processing to obtain multiple splicing sample features; Based on each of the spliced ​​sample features, a masked attention mechanism is called to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking each of the independent samples in the spliced ​​sample features respectively; Determining a loss function based on the difference between the predicted text and the sample label; The text processing model is subjected to parameter update processing based on the loss function to obtain the trained text processing model.

2. The method according to claim 1, characterized in that The step of splicing at least two of the plurality of independent sample texts to obtain a plurality of spliced ​​sample texts includes: Performing multiple selection processes on at least two of the multiple independent sample texts to obtain multiple selected sample combinations to be spliced, wherein each of the sample combinations to be spliced ​​includes at least two of the independent sample texts; The following processing is performed for each combination of samples to be spliced: Randomly combine each of the independent sample texts in the sample combination to be spliced ​​into a sample sequence; Each of the independent sample texts in the sample sequence is separated by a separator to obtain a concatenated text sample.

3. The method according to claim 1, characterized in that The step of calling the mask attention mechanism based on each of the spliced ​​sample features to perform prediction processing to obtain the predicted text includes: The following processing is performed for each of the spliced ​​sample features: Determining a key matrix, a value matrix, and a query matrix of the spliced ​​sample features; Determine a mask matrix corresponding to each of the independent sample texts in the spliced ​​sample features; Determining an attention weight value of the spliced ​​sample feature based on the key matrix, the value matrix, the query matrix, and each of the mask matrices; Based on the attention weight value and the concatenated sample feature, the next sentence text of the concatenated sample text is predicted to obtain the predicted text.

4. The method according to claim 3, characterized in that The determining the attention weight value of the spliced ​​sample feature based on the key matrix, the value matrix, the query matrix and each of the mask matrices comprises: obtaining a first product between the key matrix and the query matrix; Masking the first product based on each of the mask matrices to obtain a masked result; Obtaining a normalized result of the mask result; The product of the normalized result and the value matrix is ​​used as the attention weight value of the spliced ​​sample feature.

5. The method according to claim 3, characterized in that: The mask matrices of different types of independent sample texts are different.

6. The method according to claim 3, characterized in that When the attention weight value is represented by an attention matrix, the portion outside the mask matrix in the attention matrix is ​​a sawtooth shape, and each tooth in the sawtooth shape corresponds to a feature of an independent sample text.

7. The method according to any one of claims 1 to 6, characterized in that: The step of predicting the next sentence of the spliced ​​sample text based on the attention weight value and the spliced ​​sample feature to obtain the predicted text includes: Based on the attention weight value and the concatenated sample feature, predict multiple first prediction probabilities of each character position in the next sentence text of the concatenated sample text, wherein each of the first prediction probabilities corresponds to a candidate character; Taking the candidate character with the highest first prediction probability corresponding to each of the character positions as the target character; Each target character is combined according to the order of each character position to obtain the predicted text.

8. The method according to any one of claims 1 to 6, characterized in that: Each of the sample labels includes: a label probability sequence of the output texts corresponding to the multiple independent sample texts; The determining of a loss function based on the difference between the predicted text and the sample label includes: Obtaining a first prediction probability corresponding to each character in the predicted text; combining each of the first prediction probabilities into a prediction probability sequence; Representing the prediction probability sequence and the label probability sequence as vectors respectively; The vector space distance between the two vectors is used as the loss function.

9. The method according to any one of claims 1 to 6, characterized in that: Each of the sample labels includes: a label probability of each character in the output text corresponding to each of the multiple independent sample texts; The determining of a loss function based on the difference between the predicted text and the sample label includes: Obtaining a first prediction probability corresponding to each character in the predicted text; For each character position, the following processing is performed: Obtaining a ratio between the label probability of the character position and the first predicted probability; Obtaining a second product between the logarithm of the ratio and the label probability; The sum of the second products of each character position is used as the loss function.

10. The method according to any one of claims 1 to 6, characterized in that: After performing parameter updating processing on the text processing model based on the loss function to obtain the trained text processing model, the method further includes: In response to receiving the to-be-reply text of the target object, obtaining the historical reply text for the target object; splicing the text to be replied and the historical reply text into a spliced ​​text; The text processing model is called based on the concatenated text to perform text prediction processing to obtain a current reply text, wherein the current reply text is the reply content of the text to be replied.

11. A training device for a text processing model, characterized in that: The device comprises: A data acquisition module is configured to acquire a sample text set, wherein the sample text set includes a plurality of independent sample texts and sample labels of the plurality of independent sample texts; The data acquisition module is configured to perform splicing processing on at least two of the multiple independent sample texts to obtain multiple spliced ​​sample texts; A model training module is configured to call the initialized text processing model to perform feature extraction processing based on the multiple spliced ​​sample texts to obtain multiple spliced ​​sample features; The model training module is configured to call a masked attention mechanism based on each of the spliced ​​sample features to perform prediction processing to obtain a predicted text, wherein the masked attention mechanism includes: masking each of the independent samples in the spliced ​​sample features respectively; The model training module is configured to determine a loss function based on the difference between the predicted text and the sample label; The model training module is configured to perform parameter update processing on the text processing model based on the loss function to obtain the trained text processing model.

12. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions; A processor, used to implement the text processing model training method described in any one of claims 1 to 10 when executing the computer executable instructions or computer program stored in the memory.

13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the training method of the text processing model described in any one of claims 1 to 10 is implemented.

14. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the training method of the text processing model described in any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Sample injection method and device of model, equipment and storage medium

    CN121009369A

  • A sample injection method, apparatus, device, and storage medium for a model.

    CN121009369B

  • Fine adjustment method of retrieval reply large model and retrieval reply method and device

    CN121166898A