A text processing method, device and related equipment

By configuring adaptive components in the Bloom model and using training samples to enhance unique language features, the problem of insufficient expression ability of the Bloom model in multi-lingual environments is solved, and better text processing effect is achieved.

CN119474325BActive Publication Date: 2025-07-11ZHONGHAO XINYING (HANGZHOU) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510047891.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-07-11
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing Bloom model has poor performance in dealing with unique language features among different languages, especially in natural language processing tasks in multilingual environments.

Method used

By configuring the adapter component in the decoding module of the Bloom model, freezing the backbone model parameters, training the adapter component with training samples of different language types, obtaining and enhancing the unique language features of the pending text, including the nonlinear transformation and overlay processing of the first unique feature acquisition module, the activation module and the second unique feature acquisition module.

Benefits of technology

It improves the unique language feature expression ability of the Bloom model in different language environments and enhances the effect of multilingual text processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474325B_ABST
    Figure CN119474325B_ABST
Patent Text Reader

Abstract

The present application discloses a text processing method, apparatus and related equipment. The method includes: obtaining a text to be processed; determining the language type of the text to be processed; configuring an adaptation component for a preset language model based on the language type, where the adaptation component is pre-trained based on training samples of the language type; obtaining the target unique language features of the text to be processed based on the configured preset language model; and performing subsequent processing on the text to be processed based on the target unique language features of the text to be processed to obtain the processing result of the text to be processed. In the text processing method in the embodiments of the present application, by configuring an adaptation component in the language model, the adaptation component can be used to obtain unique language features of different languages, so that when processing texts to be processed of different language types, the expression ability of the unique language features of different languages can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a text processing method, apparatus, and related devices. Background Art

[0002] In the field of artificial intelligence, cross - language understanding is an important research direction. Currently, existing artificial intelligence models already have excellent capabilities in processing single - language tasks. However, in a multilingual environment, especially when dealing with natural language processing (NLP) tasks in multiple languages, these artificial intelligence models still face many challenges.

[0003] For example, the currently commonly used Bloom (Big Science Large Open - science Open - access Multilingual Language Model) model can represent general features between different languages well, has a certain generalization ability, and can be used across languages to a certain extent. However, in terms of features such as the unique expressions, meanings, and grammatical structures of different languages, the Bloom model still performs poorly. Summary of the Invention

[0004] The main objective of this application is to propose a text processing method, apparatus, and related devices, aiming to solve the problem that the existing Bloom model has poor performance in representing unique language features between different languages.

[0005] To achieve the above objective, this application proposes a text processing method, including:

[0006] Obtain the text to be processed;

[0007] Determine the language type of the text to be processed;

[0008] Configure an adaptation component for a preset language model based on the language type; wherein, the adaptation component is pre - trained based on training samples of this language type;

[0009] Obtain the target unique language features of the text to be processed based on the configured preset language model;

[0010] Perform subsequent processing on the text to be processed based on the target unique language features of the text to be processed to obtain the processing result of the text to be processed.

[0011] In an embodiment of this application, the adaptation component includes a first unique feature acquisition module, an activation module, and a second unique feature acquisition module that are communicatively connected in sequence. The preset model is the Bloom model. Configuring an adaptation component for a preset language model based on the language type includes:

[0012] Configure the adaptation component after each multi-layer perceptron of at least one decoding module of the Bloom model;

[0013] Configure adaptation parameters corresponding to the language type of the text to be processed for the adaptation component.

[0014] In the embodiments of the present application, the adaptation component is pre-trained based on the following method:

[0015] After configuring the adaptation component behind the Bloom model, freeze the backbone model parameters of the Bloom model, and train the adaptation component respectively based on training samples of different language types to obtain multiple sets of adaptation parameters; among them, each set of adaptation parameters corresponds to a language type respectively.

[0016] In the embodiments of the present application, obtaining the target unique language feature of the text to be processed based on the configured preset language model includes:

[0017] Based on the first unique feature acquisition module, receive the original text feature output by the multi-layer perceptron, and obtain the first unique language feature in the original text feature;

[0018] Based on the activation module, perform a non-linear transformation on the first unique language feature to obtain a transformation result;

[0019] Based on the second unique feature acquisition module, obtain the second unique language feature in the transformation result;

[0020] Use the second unique language feature as the target unique language feature of the text to be processed.

[0021] In the embodiments of the present application, performing subsequent processing on the text to be processed based on the target unique language feature of the text to be processed to obtain the processing result of the text to be processed includes:

[0022] Based on the second unique feature acquisition module, superimpose the target unique language feature on the original text feature to obtain an enhanced feature;

[0023] Based on the enhanced feature, perform subsequent processing on the text to be processed to obtain the processing result of the text to be processed.

[0024] The present application also proposes a text processing device, and the text processing device includes:

[0025] A text acquisition module, configured to acquire the text to be processed;

[0026] A processing module, configured to determine the language type of the text to be processed;

[0027] Configure an adaptation component for the preset language model based on the language type, where the adaptation component is pre-trained based on training samples of the language type;

[0028] Obtain the target unique language feature of the text to be processed based on the configured preset language model;

[0029] Perform subsequent processing on the text to be processed based on the target unique language feature of the text to be processed to obtain the processing result of the text to be processed.

[0030] In the embodiment of the present application, the adaptation component includes a first unique feature acquisition module, an activation module, and a second unique feature acquisition module that are communicatively connected in sequence. The preset model is the Bloom model, and the processing module is configured to configure the adaptation component for the preset language model based on the following method based on the language type:

[0031] Configure the adaptation component after each multi-layer perceptron of at least one decoding module of the Bloom model;

[0032] Configure adaptation parameters corresponding to the language type of the text to be processed for the adaptation component.

[0033] In the embodiment of the present application, the adaptation component is pre-trained based on the following method:

[0034] After configuring the adaptation component behind the Bloom model, freeze the backbone model parameters of the Bloom model, and train the adaptation component respectively based on training samples of different language types to obtain multiple sets of adaptation parameters; where each set of adaptation parameters corresponds to a language type respectively.

[0035] In the embodiment of the present application, the processing module is further configured to:

[0036] Receive the original text feature output by the multi-layer perceptron based on the first unique feature acquisition module, and obtain the first unique language feature in the original text feature;

[0037] Perform a non-linear transformation on the first unique language feature based on the activation module to obtain a transformation result;

[0038] Obtain the second unique language feature in the transformation result based on the second unique feature acquisition module;

[0039] Use the second unique language feature as the target unique language feature of the text to be processed.

[0040] In the embodiment of the present application, the processing module is further configured to:

[0041] Based on the second unique feature acquisition module, the target unique language feature is superimposed on the original text feature to obtain an enhanced feature;

[0042] Based on the enhanced feature, subsequent processing is performed on the text to be processed to obtain a processing result of the text to be processed.

[0043] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the text processing method described in any one of the above is implemented.

[0044] An embodiment of the present application also provides a computing device, including a processor, on which a computer program is stored, and when the computer program is executed, the text processing method described in any one of the above embodiments is implemented.

[0045] In the text processing method in the embodiment of the present application, by configuring an adaptation component in a language model, the adaptation component can be used to obtain unique language features of different languages, so that when processing text to be processed of different language types, the expression ability of unique language features of different languages can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0047] Figure 1 It is an architecture diagram of the Bloom model in the prior art;

[0048] Figure 2 It is an architecture diagram of the decoding module in the Bloom model in the prior art;

[0049] Figure 3 It is a step diagram of the text processing method in an embodiment of the present application;

[0050] Figure 4 It is an architecture diagram of the decoding module configured with an adaptation component in an embodiment of the present application;

[0051] Figure 5 It is a module diagram of the adaptation component in an embodiment of the present application;

[0052] Figure 6 It is a module diagram of the text processing device in an embodiment of the present application;

[0053] Figure 7Module diagram of a computer-readable storage medium in an embodiment of the present application;

[0054] Figure 8 Module diagram of a computing device in an embodiment of the present application.

[0055] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0056] Hereinafter, the principles and spirit of the present application will be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present application, and do not limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to convey the scope of the present disclosure fully to those skilled in the art.

[0057] Those skilled in the art know that the embodiments of the present application can be implemented as a system, a device, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0058] According to the embodiments of the present application, a text processing method, apparatus, and related devices are provided.

[0059] In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0060] Hereinafter, the principles and spirit of the present application will be elaborated in detail with reference to several representative embodiments of the present application.

[0061] Exemplary method

[0062] As Figure 1 、 Figure 2 shown, FIG. 1 is an architecture diagram of the Bloom model in the prior art, Figure 2 and FIG. is an architecture diagram of the decoding module in the prior art, where the Bloom model mainly includes an embedding layer (Embedding layer) and multiple decoding modules, and each decoding module includes multiple parallel multilayer perceptrons (Multilayer Perceptron, MLP) and multi-head attention modules (Multi-Head Attention).

[0063] Among them, the Embedding layer is mainly used to map the input text to be processed into word vectors; the multi-layer perceptron is used to extract and transform the word vectors output by the embedding layer; the multi-head attention module adds positional information (ALiBi) to the standard multi-head attention (HeadFusion), enabling the model to better understand the positional relationship of the text to be processed.

[0064] In addition, a normalization layer (Layer Normalization, LN) can be added to the output of each layer. LN is a commonly used regularization method in deep learning, mainly used to normalize the output of each layer to ensure that the data is within the range of mean 0 and variance 1, so that the gradient calculation result will not become too large or too small due to chain calculation. When processing sequence data, it can effectively prevent gradient explosion or gradient disappearance, thereby improving the expressive ability and training efficiency of the model.

[0065] After being trained, the Bloom model can have a good representation effect on the common features between different languages and has a certain generalization ability, enabling it to be used cross-linguistically to a certain extent. However, for the unique language features of different languages, the Bloom model has poor expressive ability. Therefore, this application proposes a text processing method to solve the problem that the existing natural language processing solutions have poor expressive ability for the unique language features of different languages.

[0066] As Figure 3 shown, this application embodiment proposes a text processing method, including the following steps S100 - S500:

[0067] Step S100: Obtain the text to be processed.

[0068] In this application embodiment, the text to be processed can be texts of different language types, which can be texts to be translated, or texts to be sentiment - analyzed, or texts to be classified, etc.

[0069] In this application embodiment, the text to be processed can be input into the Bloom model through the input port of the Bloom model that has been trained in the prior art and has a certain generalization ability, and the input port of the Bloom model is the port for users to input the text to be processed in the prior - art Bloom model.

[0070] Step S200: Determine the language type of the text to be processed.

[0071] In the embodiments of the present application, when inputting the text to be processed, a language identifier can also be input into the Bloom model, and this language identifier is used to represent the language type of the text to be processed. For example, when the text to be processed is in English, its language identifier can be "en", and when the text to be processed is in Chinese, its language identifier can be "cn", and so on. Then, when the user inputs the text to be processed and its corresponding language identifier, the language type of the text to be processed can be determined according to the language identifier.

[0072] Step S300: Configure an adaptation component for the preset language model based on the language type; wherein, the adaptation component is pre-trained based on the training samples of this language type.

[0073] In step S100, the text to be processed is input into the trained Bloom model. That is, in the embodiments of the present application, the preset language model is this trained Bloom model with a certain generalization ability. Among them, the overall architecture of the Bloom model proposed in the embodiments of the present application can be referred to Figure 1 as shown.

[0074] After determining that the preset language model is the Bloom model, in the embodiments of the present application, an adaptation component can be configured for this Bloom model based on the following steps S310 - S320:

[0075] Step S310: Configure the adaptation component after each multi-layer perceptron in at least one decoding module of the Bloom model.

[0076] As Figure 1 , Figure 2 , Figure 4 shown, the Bloom model includes multiple decoding modules, and each decoding module includes multiple multi-layer perceptrons. In the embodiments of the present application, when configuring the adaptation component, the adaptation component can be set after each multi-layer perceptron in at least one decoding module of the Bloom model. That is, for each decoding module of the Bloom model, it is not necessary to set the above adaptation component in all decoding modules, and the number of configured adaptation components can be determined based on the expression ability of the existing Bloom model for different languages. In addition, the adaptation component can be configured between the decoding module and the normalization layer after the decoding module.

[0077] For example, for a certain commonly used language, the Bloom model can already express its general features and unique language features well. In this case, there is no need to set up adaptation components in the Bloom model, or only a few decoding modules can be set up with adaptation components. For less commonly used languages, when the Bloom model has poor expression ability for their unique language features, adaptation components can be set up in most or all of the decoding modules to improve the expression ability for their unique language features. That is, the stronger the expression ability of the Bloom model for a certain language, the fewer the number of adaptation components that need to be set up when processing the text to be processed of this language type; the weaker the expression ability of the Bloom model for a certain language, the more the number of adaptation components that need to be set up when processing the text to be processed of this language type.

[0078] In addition, in the embodiments of the present application, when it is only necessary to set up adaptation components in some decoding modules, they can be set up from back to front in the order of the decoding modules, that is, priority is given to setting up adaptation components in the decoding modules with a higher position in the Bloom model. Moreover, when setting up adaptation components in a certain decoding module, they can be set up in all the multi-layer perceptrons of this decoding module.

[0079] Step S320: Configure adaptation parameters corresponding to the language type of the text to be processed for the adaptation components.

[0080] In the embodiments of the present application, the adaptation components are configured to be pre-trained based on the following method:

[0081] After configuring the adaptation components in the Bloom model, freeze the backbone model parameters of the Bloom model, and train the adaptation components respectively based on training samples of different languages to obtain multiple sets of adaptation parameters; where each set of adaptation parameters corresponds to a language type.

[0082] In the embodiments of the present application, after configuring the adaptation components in the Bloom model, the backbone model is the other modules in the Bloom model except for the adaptation components. The backbone model is mainly used to obtain the original text features of different training samples, and the adaptation components are used to enhance the unique language features in the original text features. Therefore, during training, the parameters of the backbone model can be frozen, and only the adaptation components are trained, which can greatly reduce the computational complexity during training.

[0083] In addition, during training, the adaptation components can also be trained respectively based on training samples of different languages. For example, by training the adaptation components only based on Chinese samples, the obtained adaptation parameters can better express the unique language features of Chinese. Another example is that by training the adaptation components only based on English-type training samples, the obtained adaptation parameters can better express the unique language features of English.

[0084] Therefore, the adaptation component is trained using training samples of multiple language types to obtain adaptation parameters corresponding to different languages. When facing the text to be processed in different languages, the adaptation component is configured based on the adaptation parameters corresponding to the language, that is, the unique language features of the text to be processed in that language can be well expressed. That is, in the embodiments of the present application, the Bloom model is further configured to: based on the language type of the text to be processed, configure the adaptation component with adaptation parameters corresponding to the language type of the text to be processed to process the text to be processed.

[0085] Step S400: Obtain the target unique language features of the text to be processed based on the configured preset language model.

[0086] In the embodiments of the present application, the unique language features of the text to be processed can be obtained through the following steps S410 - S430:

[0087] Step S410: Based on the first unique feature acquisition module, receive the original text features output by the multi-layer perceptron and obtain the first unique language features in the original text features.

[0088] As Figure 5 shown, in the embodiments of the present application, the adaptation component includes a first unique feature acquisition module 100, an activation module 200, and a second unique feature acquisition module 300 that are communicatively connected in sequence. When configuring the adaptation component, the first unique feature acquisition module 100, the activation module 200, and the second unique feature acquisition module 300 can be sequentially arranged after the multi-layer perceptron of the decoding module. The first unique feature acquisition module 100 can receive the output result of the multi-layer perceptron communicatively connected to it.

[0089] In the embodiments of the present application, after the text to be processed passes through the multi-layer perceptron of the Bloom model, the multi-layer perceptron can extract the original text features in the text to be processed. Some of these original text features belong to the common features between different languages, and some belong to the unique language features unique to each different language. An adaptation component is set after the multi-layer perceptron to increase the unique language features of different languages in the original text features output by the multi-layer perceptron, thereby enhancing the expression ability of the Bloom model for the unique language features of different languages.

[0090] In an embodiment of the present application, the first unique feature acquisition module 100 is configured to receive the original text features output by the multi-layer perceptron, reduce the dimension of the original text features, and acquire the first unique language features in the reduced-dimension original text features. For example, in an embodiment of the present application, the first unique feature acquisition module 100 is provided with a dimensionality reduction matrix W1 and a first bias matrix b1. After the first unique feature acquisition module 100 receives the original text features output by the multi-layer perceptron, assuming the original text features are h, the dimensionality reduction matrix is first used to reduce the dimension of the original text features. After being reduced in dimension by the dimensionality reduction matrix, it can obtain , and then it is biased by the first bias matrix b1 to obtain , that is, the first unique language features.

[0091] Among them, the dimensionality reduction matrix can reduce the dimension of the original text features, which can reduce the computational amount in the training process. After the original text features are reduced in dimension, the first bias matrix is introduced. Through the first bias matrix, the first unique language features after the dimension reduction of the original text features under different language types can be captured and enhanced, so as to improve the expression ability of the adaptation component for the unique language features of different language types.

[0092] Step S420: Based on the activation module, perform a non-linear transformation on the first unique language features to obtain a transformation result.

[0093] As Figure 5 shown, in an embodiment of the present application, the activation module 200 is configured to receive the output result of the first unique feature acquisition module 100 and perform a non-linear transformation on the output result of the first unique feature acquisition module 100 to obtain a transformation result. For example, in an embodiment of the present application, the activation module 200 is provided with a non-linear activation function, such as GeLu. Assuming represents the activation function, after the output result of the first unique feature acquisition module 100 is input into the activation module 200, the transformation result can be obtained after passing through the non-linear activation function. Among them, the first unique feature acquisition module 100 reduces the dimension of the original text features and enhances the unique language features of different languages through the first bias matrix, and then inputs it into the activation module 200 for non-linear transformation, which can further increase the expression ability of the adaptation component for the unique language features of different languages.

[0094] Step S430: Based on the second unique feature acquisition module, acquire the second unique language features in the transformation result.

[0095] ‌As Figure 5As shown, in the embodiment of the present application, the second unique feature acquisition module 300 is used to receive the transformation result output by the activation module 200, perform dimensionality increase on the transformation result, and acquire the second unique language feature in the transformation result after dimensionality increase. For example, in the embodiment of the present application, the second unique feature acquisition module 300 is provided with a dimensionality increase matrix W2 and a second bias matrix b2. After the activation module 200 outputs the transformation result to the second unique feature acquisition module 300, the second unique feature acquisition module 300 first performs dimensionality increase on the transformation result to obtain , and then performs bias through the second bias matrix to obtain , that is, the second unique language feature.

[0096] Among them, the second unique feature acquisition module 300 increases the dimensionality of the transformation result to the original dimension of the original text feature, which can facilitate subsequent operations in the Bloom model, while the second bias matrix can capture the unique language features after dimensionality increase of the original text features under different language types and enhance them, so as to further improve the expression ability of the adaptation component for the unique language features of different language types.

[0097] Step S430: Use the second unique language feature as the target unique language feature of the text to be processed.

[0098] In the embodiment of the present application, in step S420, the second unique feature acquisition module 300 can output , that is, the second unique language feature. In step S430, the second unique language feature output by the second unique feature acquisition module 300 can be used as the target unique language feature of the text to be processed.

[0099] Step S500: Perform subsequent processing on the text to be processed based on the target unique language feature of the text to be processed to obtain the processing result of the text to be processed.

[0100] In the embodiment of the present application, the target unique language feature can be superimposed on the original text feature based on the second unique feature acquisition module 300 to obtain an enhanced feature, and subsequent processing is performed on the text to be processed based on the enhanced feature.

[0101] Among them, the second unique language feature is the target unique language feature, specifically , the original text is h, and using the second unique feature acquisition module 300 to superimpose the two can enhance the unique language feature in the original text feature. Specifically, the superimposition can be performed based on the following formula:

[0102] +h

[0103] Among them, h represents the original text feature of the multi-layer perceptron. represents a dimensionality reduction matrix, represents a first bias matrix, represents an activation function, represents a dimensionality increase matrix, represents a second bias matrix, represents the enhanced feature after superposition.

[0104] In the embodiment of the present application, after the adaptation component obtains the enhanced feature after superposition, it outputs the enhanced feature to other modules of the Bloom model, which can be further processed by other subsequent modules. Since the enhanced feature has enhanced the unique language features of the text to be processed, the processing results obtained in the subsequent processing process will surely have a better expression effect on the unique language features of the text to be processed.

[0105] In the text processing method of the embodiment of the present application, by configuring an adaptation component in the language model, the adaptation component can obtain the unique language features of different languages, so as to enhance the expression ability of the unique language features of different languages when processing texts to be processed of different language types.

[0106] Exemplary device

[0107] As Figure 6 shown, this exemplary embodiment proposes a text processing device 400, and the text processing device 400 includes:

[0108] A text acquisition module 410, configured to acquire a text to be processed;

[0109] A processing module 420, configured to determine the language type of the text to be processed;

[0110] Configure an adaptation component for a preset language model based on the language type; wherein, the adaptation component is pre-trained based on training samples of this language type;

[0111] Obtain the target unique language features of the text to be processed based on the configured preset language model;

[0112] Perform subsequent processing on the text to be processed based on the target unique language features of the text to be processed to obtain a processing result of the text to be processed.

[0113] In the embodiment of the present application, the adaptation component includes a first unique feature acquisition module 100, an activation module 200, and a second unique feature acquisition module 300 that are communicatively connected in sequence. The preset model is a Bloom model, and the processing module 420 is configured to configure an adaptation component for a preset language model based on the following method based on the language type:

[0114] Configure the adaptation component after each multi-layer perceptron of at least one decoding module of the Bloom model;

[0115] Configure adaptation parameters corresponding to the language type of the text to be processed for the adaptation component.

[0116] In the embodiments of the present application, the adaptation component is pre-trained based on the following method:

[0117] After configuring the adaptation component behind the Bloom model, freeze the backbone model parameters of the Bloom model, and train the adaptation component respectively based on training samples of different language types to obtain multiple groups of adaptation parameters; where each group of adaptation parameters corresponds to a language type respectively.

[0118] In the embodiments of the present application, the processing module 420 is further configured to:

[0119] Receive the original text features output by the multi-layer perceptron based on the first unique feature acquisition module 100, and obtain the first unique language feature in the original text features;

[0120] Perform a non-linear transformation on the first unique language feature based on the activation module 200 to obtain a transformation result;

[0121] Obtain the second unique language feature in the transformation result based on the second unique feature acquisition module 300;

[0122] Use the second unique language feature as the target unique language feature of the text to be processed.

[0123] In the embodiments of the present application, the processing module 420 is further configured to:

[0124] Superimpose the target unique language feature on the original text feature based on the second unique feature acquisition module 300 to obtain an enhanced feature;

[0125] Perform subsequent processing on the text to be processed based on the enhanced feature to obtain the processing result of the text to be processed.

[0126] For the specific processing methods of each module of the text processing device 400 in the embodiments of the present application, reference can be made to each embodiment of the text processing method in the above exemplary method, which will not be elaborated here one by one.

[0127] The text processing device 400 in the embodiments of the present application configures an adaptation component in the language model through the processing module 420, and the adaptation component can obtain unique language features of different languages, so that when processing texts to be processed of different language types, the expression ability of the unique language features of different languages can be enhanced.

[0128] Exemplary medium

[0129] After introducing the methods, media, and systems of the exemplary embodiments of the present application, next, reference is made to Figure 7 to describe the computer-readable storage media of the exemplary embodiments of the present application. Please refer to Figure 7 , which shows that the computer-readable storage medium is an optical disc 70, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will implement the various steps recorded in the above method embodiments. For example, obtaining the text to be processed; determining the language type of the text to be processed; configuring an adaptation component for a preset language model based on the language type; wherein, the adaptation component is pre-trained based on the training samples of this language type; obtaining the target unique language features of the text to be processed based on the configured preset language model; performing subsequent processing on the text to be processed based on the target unique language features of the text to be processed to obtain the processing result of the text to be processed. The specific implementation manners of each step will not be repeated here. It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated one by one here.

[0130] Exemplary computing device

[0131] After introducing the methods, systems, and media of the exemplary embodiments of the present application, next, reference is made to Figure 8 the computing devices of the exemplary embodiments of the present application.

[0132] Figure 8 The block diagram of an exemplary computing device 80 suitable for implementing the embodiments of the present application is shown. The computing device 80 may be a computer system or a server. Figure 8 The shown computing device 80 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0133] As Figure 8 shown, the components of the computing device 80 may include, but are not limited to: one or more processors or processing units 801, a system memory 802, and a bus 803 connecting different system components (including the system memory 802 and the processing unit 801).

[0134] Computing device 80 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computing device 80, including volatile and non-volatile media, removable and non-removable media.

[0135] System memory 802 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 8021 and / or cache memory 8022. Computing device 80 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 8023 can be used to read and write to non-removable, non-volatile magnetic media ( Figure 8 not shown in the figure, commonly referred to as a "hard disk drive"). Although not shown in Figure 8 the figure, a disk drive can be provided for reading and writing to a removable non-volatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical media). In these cases, each drive can be connected to bus 803 through one or more data media interfaces. System memory 802 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present application.

[0136] A program / utility 8025 having a set (at least one) of program modules 8024 can be stored, for example, in system memory 802, and such program modules 8024 include but are not limited to: an operating system, one or more application programs, other program modules, and program data, and the implementation of a network environment may be included in each or some combination of these examples. Program modules 8024 generally perform the functions and / or methods in the embodiments described in the present application.

[0137] Computing device 80 can also communicate with one or more external devices 804 (such as a keyboard, pointing device, display, etc.). Such communication can be carried out through an input / output (I / O) interface. Moreover, computing device 80 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through network adapter 806. As Figure 8 shown, network adapter 806 communicates with other modules (such as processing unit 801, etc.) of computing device 80 through bus 803. It should be understood that although Figure 8 not shown in the figure, other hardware and / or software modules can be used in conjunction with computing device 80.

[0138] The processing unit 801 executes various functional applications and data processing by running programs stored in the system memory 802. For example, it obtains the text to be processed, determines the language type of the text to be processed, configures an adaptation component for a preset language model based on the language type, where the adaptation component is pre-trained based on training samples of this language type, obtains the target unique language features of the text to be processed based on the configured preset language model, and performs subsequent processing on the text to be processed based on the target unique language features of the text to be processed to obtain the processing result of the text to be processed. The specific implementation manners of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the computing device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0139] In the description of the present application, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0140] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0141] In several embodiments provided in the present application, it should be understood that the disclosed systems, systems, and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of systems or units can be in electrical, mechanical, or other forms.

[0142] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0143] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0144] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computing device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0145] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0146] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

Claims

1. A text processing method, the text processing method comprising: Obtaining the text to be processed; Determining the language type of the text to be processed; Configuring an adaptation component for a preset language model based on the language type; wherein, the adaptation component is pre-trained based on training samples of this language type, and the adaptation parameters of the adaptation component are only pre-trained based on training samples of this language type; Obtaining the target unique language feature of the text to be processed based on the configured preset language model; Performing subsequent processing on the text to be processed based on the target unique language feature of the text to be processed to obtain the processing result of the text to be processed; The adaptation component includes a first unique feature acquisition module, an activation module, and a second unique feature acquisition module that are communicatively connected in sequence. The preset model is the Bloom model. Configuring the adaptation component for the preset language model based on the language type includes: Configuring the adaptation component after each multi-layer perceptron of at least one decoding module of the Bloom model; Configuring adaptation parameters corresponding to the language type of the text to be processed for the adaptation component; The obtaining the target unique language feature of the text to be processed based on the configured preset language model includes: Based on the first unique feature acquisition module receiving the original text feature output by the multi-layer perceptron and obtaining the first unique language feature in the original text feature; Performing a non-linear transformation on the first unique language feature based on the activation module to obtain a transformation result; Based on the second unique feature acquisition module obtaining the second unique language feature in the transformation result; Taking the second unique language feature as the target unique language feature of the text to be processed.

2. The text processing method according to claim 1, wherein, The adaptation component is pre-trained based on the following method: After configuring the adaptation component behind the Bloom model, freezing the backbone model parameters of the Bloom model, and respectively training the adaptation component based on training samples of different language types to obtain multiple groups of adaptation parameters; wherein, each group of adaptation parameters corresponds to a language type.

3. The text processing method according to claim 1, wherein, The performing subsequent processing on the text to be processed based on the target unique language feature of the text to be processed to obtain the processing result of the text to be processed includes: Based on the second unique feature acquisition module, superimposing the target unique language feature on the original text feature to obtain an enhanced feature; Performing subsequent processing on the text to be processed based on the enhanced feature to obtain the processing result of the text to be processed.

4. A text processing device, the text processing device comprising: A text acquisition module, configured to obtain the text to be processed; A processing module, configured to determine the language type of the text to be processed; Configuring an adaptation component for a preset language model based on the language type; wherein, the adaptation component is pre-trained based on training samples of this language type; Obtaining the target unique language feature of the text to be processed based on the configured preset language model; Performing subsequent processing on the text to be processed based on the target unique language feature of the text to be processed to obtain the processing result of the text to be processed; The adaptation component includes a first unique feature acquisition module, an activation module, and a second unique feature acquisition module that are communicatively connected in sequence. The preset model is the Bloom model. The processing module is configured to configure the adaptation component for the preset language model based on the following method according to the language type: Configure the adaptation component after each multi-layer perceptron of at least one decoding module of the Bloom model; Configure adaptation parameters corresponding to the language type of the text to be processed for the adaptation component; The processing module is further configured to: Receive the original text features output by the multi-layer perceptron based on the first unique feature acquisition module, and obtain the first unique language feature in the original text features; Perform a non-linear transformation on the first unique language feature based on the activation module to obtain a transformation result; Obtain the second unique language feature in the transformation result based on the second unique feature acquisition module; Use the second unique language feature as the target unique language feature of the text to be processed.

5. The text processing device according to claim 4, wherein The adaptation component is pre-trained based on the following method: After configuring the adaptation component behind the Bloom model, freeze the backbone model parameters of the Bloom model, and train the adaptation component respectively based on training samples of different language types to obtain multiple groups of adaptation parameters; where each group of adaptation parameters corresponds to a language type respectively.

6. The text processing device according to claim 4, wherein, The processing module is further configured to: Superimpose the target unique language feature on the original text feature based on the second unique feature acquisition module to obtain an enhanced feature; Perform subsequent processing on the text to be processed based on the enhanced feature to obtain the processing result of the text to be processed.

7. A computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the method according to claims 1-3.

8. A computing device, including a processor, and a computer program is stored on the processor, and when the computer program is executed, it implements the processing method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Speech recognition model training method and device, speech recognition method and device, equipment and medium

    CN117711386A

  • Text processing method, product, equipment and medium

    CN118396126A