Text processing method and device and related equipment
By determining features at word level and sentence level and combining contextual representation and attention weight, the text processing method is solved, and the problem of large converter models consume large computing resources and low translation accuracy when processing long text is solved, achieving more efficient calculations and more accurate translation.
Patent Information
- Application Number
- CN202510474040.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the existing converter model processes text with longer lengths, computing resources are consumed more, and the global self-attention mechanism leads to low translation accuracy.
A text processing method is proposed to generate text features by determining word features and sentence features at word level and sentence level, combining context representation and attention weight. This method adopts bidirectional RNN neural network and semantic spatial projection technology to process text progressively and reduce the calculation amount of the global self-attention mechanism.
It effectively reduces the length of the input sequence of the converter large model, reduces computing resource consumption, and improves translation accuracy, avoiding the problem of weakening of local dependencies.
Smart Images

Figure CN119990103A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a text processing method, apparatus and related equipment. Background Art
[0002] With the acceleration of globalization, the demand for machine translation systems is growing. Current machine translation systems often introduce large transformer models. The transformer model is a model that relies on the self-attention mechanism, which allows the model to process sequence data without relying on the traditional recurrent neural network structure, thereby significantly improving processing speed and efficiency.
[0003] However, there are still many limitations when using the existing large converter model for machine translation. For example, when faced with long texts, the standard converter model uses a global self-attention mechanism, which consumes a lot of computing resources. In addition, the global self-attention mechanism also weakens local dependencies, which in turn affects translation accuracy. Summary of the invention
[0004] The main purpose of this application is to propose a text processing method, device and related equipment, aiming to solve the problem that the large converter model in the prior art consumes large computing resources and has low translation accuracy when facing long texts.
[0005] To achieve the above purpose, this application proposes a text processing method, including: Get the text to be processed; Determine each sentence contained in the text to be processed, and the words contained in each sentence; Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence; Based on the sentence features of each sentence, determine the context representation of each sentence; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
[0006] In the embodiment of the present application, the sentence features of each sentence are determined based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence, including: Based on the first preset neural network, determining the word features of each word contained in each sentence respectively; Based on the projection of each word’s feature in the semantic space, determine the attention weight of the word feature corresponding to each word in the sentence where the word is located; Based on the word features of each word contained in each sentence and the corresponding attention weights, the sentence features of each sentence are determined.
[0007] In the embodiment of the present application, determining the context representation of each sentence based on the sentence features of each sentence includes: Based on the second preset neural network and the sentence features of each sentence, the context representation of each sentence is determined.
[0008] In the embodiment of the present application, the text features of the text to be processed are determined based on the context representation of each sentence and the attention weight of each sentence, including: Determine the attention weight of each sentence based on the projection of its sentence features in the semantic space; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
[0009] The present application also proposes a text processing device, the text processing device comprising: An acquisition module is used to obtain the text to be processed; A processing module, used for determining each sentence contained in the text to be processed, and words contained in each sentence; Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence; Based on the sentence features of each sentence, determine the context representation of each sentence; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
[0010] In the embodiment of the present application, the processing module is also used for: Based on the first preset neural network, determining the word features of each word contained in each sentence respectively; Based on the projection of each word’s feature in the semantic space, determine the attention weight of the word feature corresponding to each word in the sentence where the word is located; Based on the word features of each word contained in each sentence and the corresponding attention weights, the sentence features of each sentence are determined.
[0011] In the embodiment of the present application, the processing module is also used for: Based on the second preset neural network and the sentence features of each sentence, the context representation of each sentence is determined.
[0012] In the embodiment of the present application, the processing module is also used for: Determine the attention weight of each sentence based on the projection of its sentence features in the semantic space; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
[0013] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the text processing method described in any of the above embodiments is implemented.
[0014] An embodiment of the present application further provides a computing device, including a processor, on which a computer program is stored, and when the computer program is executed, the text processing method described in any of the above embodiments is implemented.
[0015] The text processing method in the embodiment of the present application first determines the word features, projections of word features, and attention weights of each word in each sentence at the word level to obtain sentence features, and then obtains text features at the text level based on the sentence features obtained at the word level, and obtains text features at the text level based on the context representation, sentence features, projections of sentence features, and attention weights. The text features obtained based on this layer-by-layer progressive processing method from word to sentence to text can not only represent the text to be processed globally, but also fully consider the local features of each sentence and each word contained in the text to be processed, which can avoid weakening local dependencies and improve processing accuracy. In addition, when processing the text to be processed, the text to be processed can be pre-processed based on the text processing method in the embodiment of the present application to obtain the corresponding text features, and then input them into the converter large model for processing, which can effectively reduce the length of the sequence input to the converter large model, thereby reducing the amount of calculation of the converter large model, thereby reducing the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0017] Figure 1 is a step diagram of a text processing method in an embodiment of the present application; Figure 2 is a module diagram of a text processing device in an embodiment of the present application; Figure 3 A module diagram of a computer-readable storage medium in an embodiment of the present application; Figure 4 A module diagram of a computing device in one embodiment of the present application.
[0018] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0019] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present application, and are not intended to limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0020] Those skilled in the art know that the embodiments of the present application can be implemented as a system, device, method or computer program product. Therefore, the present application can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0021] According to the implementation manner of the present application, a text processing method, apparatus and related equipment are proposed.
[0022] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.
[0023] The principles and spirit of the present application are explained in detail below with reference to several representative implementations of the present application.
[0024] Exemplary Methods like Figure 1 As shown, the embodiment of the present application proposes a text processing method, comprising the following steps S100-S500: Step S100: Obtain the text to be processed.
[0025] In the embodiment of the present application, the text to be processed may be a text to be translated, and the text to be translated may be in various language types. In addition, the text to be translated may be a long text containing multiple sentences.
[0026] Step S200: Determine each sentence contained in the text to be processed, and each word contained in each sentence.
[0027] In the application embodiment, after the text to be processed is obtained, sentences and words can be divided based on the text to be processed.
[0028] In the embodiment of the present application, a longer text to be processed can be divided into multiple shorter sentences according to punctuation marks.
[0029] For example, suppose the text to be processed is: "Today is sunny and windless, a good day for traveling. Why not call a friend and ask him to have a picnic?" According to punctuation marks, it can be divided into the following sentences: Sentence 1: It is a sunny day today; Sentence 2: There was not a breath of wind; Sentence 3: It’s a good day for traveling; Sentence 4: Why not call your friend and ask him to go on a picnic?
[0030] After dividing the text to be processed into a plurality of sentences according to punctuation marks, the words contained in each sentence obtained by the division are determined.
[0031] For example, sentence 1 in the above example includes the words: today, weather, sunny; Sentence 2 includes the words: no, a trace, wind; Sentence 3 includes the words: is, a, suitable for, travel, of, good, day.
[0032] Sentence 4 includes the words: why not, give, friend, call, make an appointment with, him, picnic, bar.
[0033] In the embodiment of the present application, the sentences and words obtained by dividing the text to be processed can also be expressed based on the following formulas (1) and (2): (1) (2) Among them, X represents the text to be processed. to Represents each sentence in the text to be processed X, represents the i-th sentence, to Representative sentences The words contained, i∈[1,n], j∈[1,m].
[0034] Step S300: Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence.
[0035] In the embodiment of the present application, the sentence features of each sentence can be determined based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence through the following steps S310-S330: Step S310: Based on the first preset neural network, determine the word features of each word contained in each sentence; Step S320: Based on the projection of the word feature of each word in the semantic space, determine the attention weight of the word feature corresponding to each word in the sentence where the word is located; Step S330: Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the corresponding attention weights.
[0036] In the embodiment of the present application, in step S310, the first preset neural network may adopt a bidirectional RNN neural network. When determining the word features of the words contained in each sentence, each word contained in the sentence may be input into the bidirectional RNN neural network, and the word features of each word may be obtained by encoding.
[0037] In the embodiment of the present application, the word feature of each word can be determined by the following formula (3): (3) in, Expressing sentences The jth word contained in represents a bidirectional RNN neural network, Representative words Word features, i∈[1,n], j∈[1,m].
[0038] The word feature of each word can be determined through step S310. In step S320, the attention weight of each word in the sentence can be determined based on the projection of the word feature of each word in the semantic space.
[0039] In the embodiment of the present application, the projection of each word in the semantic space can be determined by the following formula (4): (4) in, Representative word features, Representative word features The projection in the semantic space, i∈[1,n],j∈[1,m], and Represents calculation parameters.
[0040] In the embodiment of the present application, the semantic space represents a higher-dimensional space composed of multiple semantically related features, wherein the semantic space may include the length feature of each word, the semantic features in different contexts, the tone features, etc. Based on the above formula (4), each word feature can be modified and transformed at the word level, highlighting more important features and weakening irrelevant details, so as to obtain a more suitable representation of each word feature. In addition, the calculation parameters and Can be pre-trained.
[0041] After obtaining the projection of each word feature, the attention weight of each word in its sentence can be determined based on the following formula (5):
[0042] in, represent Corresponding sentences The query vector is Representative projection Corresponding word features In the sentence where The attention weights in , i∈[1,n], j∈[1,m].
[0043] After obtaining the attention weight of each word in its sentence, the sentence features of each sentence can be determined based on the word features of each word contained in each sentence and the corresponding attention weight.
[0044] In the embodiment of the present application, the sentence feature of each sentence can be determined based on the following formula (6): (6) in, Representative sentences Sentence features, i∈[1,n], j∈[1,m].
[0045] Through step S300, firstly, based on the bidirectional RNN neural network, the word features of each word in the sentence are obtained at the word level, and then based on the projection of the word features in the high-dimensional semantic space, each word feature is optimized, and then according to the projection of each word feature, the importance of each word feature in the sentence, that is, the attention weight, is determined, and then according to the attention weight of each word feature, the word features of each word are weighted, and then the sentence features of each sentence at the sentence level obtained based on the weighted word features can be more accurate.
[0046] Step S400: Determine the context representation of each sentence based on the sentence features of each sentence.
[0047] In an embodiment of the present application, the context representation of each sentence can be determined based on the second preset neural network and the sentence features of each sentence.
[0048] The second preset neural network can be a bidirectional RNN neural network. When obtaining the context representation of each sentence, the sentence features corresponding to each sentence can be input into the second preset neural network, encoded and the corresponding context representation can be obtained. It should be noted that the first preset neural network and the second preset neural network are both bidirectional RNN neural networks, the two have the same architecture, both are pre-trained, and the internal parameters after training are different.
[0049] In the embodiment of the present application, the context representation of each sentence can be determined based on the following formula (7): (7) in, Representative sentences The sentence features of Representative sentences Context representation, BiRNN represents the second preset neural network, i∈[1,n].
[0050] In the embodiment of the present application, the contextual relationship of each sentence, that is, the contextual representation of each sentence, can be obtained by encoding at the sentence level through the second preset neural network.
[0051] Step S500: Determine the text features of the text to be processed based on the context representation of each sentence and the attention weight of each sentence.
[0052] In the embodiment of the present application, the text features of the text to be processed can be determined based on the context representation of each sentence and the attention weight of each sentence through the following steps S510-S520: Step S510: Determine the attention weight of each sentence based on the projection of the sentence features of each sentence in the semantic space.
[0053] Step S520: Determine the text features of the text to be processed based on the context representation of each sentence and the attention weight of each sentence.
[0054] In step S510, the attention weight of each sentence can be determined based on the projection of the sentence features of each sentence in the semantic space by the following formula (8): (8) in, Representative sentences The sentence features of Representative sentence features The projection in the semantic space, i∈[1,n],j∈[1,m], and Represents calculation parameters.
[0055] In the embodiment of the present application, the semantic space represents a higher-dimensional space composed of multiple features. Based on the above formula (8), each sentence feature can be modified and transformed at the sentence level to highlight more important features and weaken irrelevant details, thereby obtaining a more accurate representation of each sentence feature. In addition, the calculation parameters and Can be pre-trained.
[0056] After obtaining the projection of each sentence feature, the attention weight of each sentence can be determined based on the following formula (9): (9) in, represent Corresponding sentences The query vector is Representative sentences The attention weight, Representative sentences Sentence features Projection in semantic space, i∈[1,n].
[0057] After obtaining the attention weight of each sentence, the text features of the text to be processed can be determined based on the sentence features of each sentence in the text to be processed and the corresponding attention weights.
[0058] In the embodiment of the present application, the text features of the text to be processed can be determined based on the following formula (10): (10) in, Representative sentences The attention weight, Representative sentences d represents the sentence features of the text to be processed X, i∈[1,n].
[0059] In steps S400 and S500, the sentence features of each sentence are first obtained at the sentence level through a bidirectional RNN neural network, and then each sentence feature is optimized based on the projection in a high-dimensional semantic space. Then, based on the projection of the sentence features, the importance of each sentence feature, i.e., the attention weight, is determined. Then, based on the attention weight of each sentence feature, the sentence features of each sentence are weighted, and then the text features at the text level obtained based on the weighted sentence features can be more accurate.
[0060] The text processing method in the embodiment of the present application determines the word features of each word in each sentence at the word level during the first encoding, and obtains the sentence features at the sentence level based on the word features, the projection of the word features, and the attention weights. During the second encoding, the context representation at the sentence level is obtained based on the sentence features obtained at the word level, and the text features at the text level are obtained based on the sentence features, the projection of the sentence features, and the attention weights. The text features obtained based on this layer-by-layer progressive processing method from word to sentence to text can not only represent the text to be processed globally, but also fully consider the local features of each sentence and each word contained in the text to be processed, which can avoid weakening of local dependencies and improve processing accuracy.
[0061] In addition, in the prior art, when processing ultra-long texts, if a large transformer model based on Transformer is used directly, its global self-attention will result in a large amount of calculation and a high dependence on computing resources. When the hierarchical attention mechanism (word level and sentence level) in the embodiment of the present application is adopted, the local context is compressed with two layers of attention, first from words to sentences, and then from sentences to texts, and finally a shorter global representation (text feature) can be obtained. Moreover, the shorter text feature can also retain global semantic information. Therefore, when processing the text to be processed, such as when translating the text to be processed, the text features of the text to be processed can be obtained based on the text processing method in the embodiment of the present application, and then input into the large transformer model based on Transformer for translation, which can effectively reduce the length of the sequence input to the large transformer model, thereby reducing the amount of calculation of the large transformer model and reducing the consumption of computing resources.
[0062] Exemplary Devices like Figure 2 As shown, this exemplary embodiment provides a text processing device 100, and the text processing device 100 includes: An acquisition module 110 is used to acquire the text to be processed; The processing module 120 is used to determine each sentence contained in the text to be processed and the words contained in each sentence; Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence; Based on the sentence features of each sentence, determine the context representation of each sentence; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
[0063] In the embodiment of the present application, the processing module 120 is further used for: Based on the first preset neural network, determining the word features of each word contained in each sentence respectively; Based on the projection of each word’s feature in the semantic space, determine the attention weight of the word feature corresponding to each word in the sentence where the word is located; Based on the word features of each word contained in each sentence and the corresponding attention weights, the sentence features of each sentence are determined.
[0064] In the embodiment of the present application, the processing module 120 is further used for: Based on the second preset neural network and the sentence features of each sentence, the context representation of each sentence is determined.
[0065] In the embodiment of the present application, the processing module is also used for: Determine the attention weight of each sentence based on the projection of its sentence features in the semantic space; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
[0066] The text processing device 100 in the embodiment of the present application, when processing the text to be processed, first determines the word features, projections of the word features, and attention weights of each word in each sentence at the word level through the processing module 120 to obtain sentence features, and then obtains the sentence features at the word level based on the context representation, sentence features, projections of sentence features, and attention weights at the sentence level to obtain text features at the text level. The text features obtained based on this layer-by-layer progressive processing method from word to sentence to text can not only represent the text to be processed globally, but also fully consider the local features of each sentence and each word contained in the text to be processed, which can avoid weakening of local dependencies and improve processing accuracy.
[0067] In addition, the text processing device 100 in the embodiment of the present application adopts a hierarchical attention mechanism (word level and sentence level), first from words to sentences, and then from sentences to text, and compresses the local context with two layers of attention, and finally a shorter global representation (text feature) can be obtained. Moreover, the shorter text feature can also retain the global semantic information. Therefore, when processing the text to be processed, such as when translating the text to be processed, the text to be processed can be preprocessed based on the text processing device 100 in the embodiment of the present application to obtain the text features of the text to be processed, and then input it into the transformer-based converter model for translation, which can effectively reduce the length of the sequence input to the converter model, thereby reducing the amount of calculation of the converter model and reducing the consumption of computing resources.
[0068] Exemplary Media After introducing the method, medium and system of the exemplary embodiments of the present application, next, reference is made to Figure 3 For a description of the computer-readable storage medium of the exemplary embodiment of the present application, please refer to Figure 3, the computer-readable storage medium shown is a CD 70, on which a computer program (i.e., a program product) is stored. When the computer program is executed by the processor, the steps recorded in the above method implementation are implemented, for example, obtaining the text to be processed; determining the sentences contained in the text to be processed, and the words contained in each sentence; determining the sentence features of each sentence based on the word features of each word contained in each sentence, and the attention weight of each word feature in its corresponding sentence; determining the context representation of each sentence based on the sentence features of each sentence; determining the text features of the text to be processed based on the context representation of each sentence and the attention weight of each sentence. The specific implementation methods of each step are not repeated here. It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not described one by one here.
[0069] Exemplary Computing Devices After introducing the methods, systems, and media of the exemplary embodiments of the present application, reference is now made to Figure 4 A computing device according to an exemplary embodiment of the present application.
[0070] Figure 4 A block diagram of an exemplary computing device 80 suitable for implementing embodiments of the present application is shown. The computing device 80 may be a computer system or a server. Figure 4 The computing device 80 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0071] like Figure 4 As shown, the components of the computing device 80 may include, but are not limited to: one or more processors or processing units 801 , a system memory 802 , and a bus 803 connecting various system components (including the system memory 802 and the processing unit 801 ).
[0072] The computing device 80 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computing device 80, including volatile and non-volatile media, removable and non-removable media.
[0073] The system storage 802 may include computer system readable media in the form of volatile storage, such as random access memory (RAM) 8021 and / or cache memory 8022. The computing device 80 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the ROM 8023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 4 is not shown in the Figure 4 As shown in FIG. 8 , a disk drive for reading and writing a removable non-volatile disk (e.g., a “floppy disk”) and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to the bus 803 via one or more data medium interfaces. The system storage 802 may include at least one program product, which has a set (e.g., at least one) of program modules, which are configured to perform the functions of each embodiment of the present application.
[0074] A program / utility 8025 having a set (at least one) of program modules 8024 may be stored, for example, in system memory 802, and such program modules 8024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 8024 generally perform the functions and / or methods of the embodiments described herein.
[0075] The computing device 80 may also communicate with one or more external devices 804 (e.g., a keyboard, a pointing device, a display, etc.). Such communication may be performed via an I / O interface 805. In addition, the computing device 80 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 806. Figure 4 As shown, the network adapter 806 communicates with other modules (such as the processing unit 801, etc.) of the computing device 80 via the bus 803. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with computing device 80 .
[0076] The processing unit 801 executes various functional applications and data processing by running the program stored in the system storage 802, for example, obtaining the text to be processed; determining the sentences contained in the text to be processed, and the words contained in each sentence; determining the sentence features of each sentence based on the word features of each word contained in each sentence, and the attention weight of each word feature in its corresponding sentence; determining the context representation of each sentence based on the sentence features of each sentence; determining the text features of the text to be processed based on the context representation of each sentence and the attention weight of each sentence. The specific implementation of each step is not repeated here. It should be noted that although several units / modules or subunits / submodules of the computing device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiment of the present application, the features and functions of two or more units / modules described above can be concretized in one unit / module. On the contrary, the features and functions of one unit / module described above can be further divided into multiple units / modules to be concretized.
[0077] In the description of the present application, it should be noted that the terms "first", "second" and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0078] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0079] In the several embodiments provided in the present application, it should be understood that the disclosed systems, systems and methods can be implemented in other ways. The system embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the system or unit can be electrical, mechanical or other forms.
[0080] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0081] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0082] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that can be executed by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computing device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.
[0083] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0084] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
Claims
1. A text processing method, the text processing method comprising: Get the text to be processed; Determine each sentence contained in the text to be processed, and the words contained in each sentence; Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence; Based on the sentence features of each sentence, determine the context representation of each sentence; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
2. The text processing method according to claim 1, wherein the sentence features of each sentence are determined based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence, comprising: Based on the first preset neural network, determining the word features of each word contained in each sentence respectively; Based on the projection of each word’s feature in the semantic space, determine the attention weight of the word feature corresponding to each word in the sentence where the word is located; Based on the word features of each word contained in each sentence and the corresponding attention weights, the sentence features of each sentence are determined.
3. The text processing method according to claim 1, wherein determining the context representation of each sentence based on the sentence features of each sentence comprises: Based on the second preset neural network and the sentence features of each sentence, the context representation of each sentence is determined.
4. The text processing method according to claim 1, wherein determining the text features of the text to be processed based on the context representation of each sentence and the attention weight of each sentence comprises: Determine the attention weight of each sentence based on the projection of its sentence features in the semantic space; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
5. A text processing device, comprising: An acquisition module is used to obtain the text to be processed; A processing module, used for determining each sentence contained in the text to be processed, and words contained in each sentence; Determine the sentence features of each sentence based on the word features of each word contained in each sentence and the attention weight of each word feature in its corresponding sentence; Based on the sentence features of each sentence, determine the context representation of each sentence; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
6. The text processing device according to claim 5, wherein the processing module is further configured to: Based on the first preset neural network, determining the word features of each word contained in each sentence respectively; Based on the projection of each word’s feature in the semantic space, determine the attention weight of the word feature corresponding to each word in the sentence where the word is located; Based on the word features of each word contained in each sentence and the corresponding attention weights, the sentence features of each sentence are determined.
7. The text processing device according to claim 5, wherein the processing module is further configured to: Based on the second preset neural network and the sentence features of each sentence, the context representation of each sentence is determined.
8. The text processing device according to claim 5, wherein the processing module is further configured to: Determine the attention weight of each sentence based on the projection of its sentence features in the semantic space; Based on the context representation of each sentence and the attention weight of each sentence, the text features of the text to be processed are determined.
9. A computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute the text processing method according to any one of claims 1 to 4.
10. A computing device, comprising a processor, wherein a computer program is stored on the processor, and when the computer program is executed, the text processing method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Neural machine translation method based on PoS (Part of Speech) attention mechanism
CN107590138A
Cross-border ethnic text classification method and device fusing domain knowledge graph
CN113901228A