Text processing method and device, computer equipment and computer readable storage medium
By aligning the text information to be processed and associated image information is performed, target alignment information is generated and decoding is performed, the difficulty of inter-language translation lacking parallel corpus is solved, and high-quality mapping from low-resource language to target language is achieved.
Patent Information
- Application Number
- CN202311446372.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
Existing machine translation techniques have difficulties in translating between languages lacking parallel corpus, especially in the case of imbalance in national and cultural development.
By obtaining the pending text information and its associated image information, an alignment operation is performed to generate the target alignment information, and the target text information is obtained through decoding processing, thereby realizing the mapping of the low-resource language to the target language.
This method reduces the noise introduced by irrelevant information, improves the processing quality of text information, and can easily achieve translation from low-resource language to target language.
Smart Images

Figure CN119940375A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to a text processing method, apparatus, computer device and computer-readable storage medium. Background Art
[0002] Traditional machine translation generally relies on a large amount of parallel bilingual corpora to translate sentences. This translation method is widely used in application scenarios such as popular science video subtitle translation and radio subtitle translation. However, in actual scenarios, due to the imbalance in development between countries and cultures, most languages in the world lack parallel corpora. The current translation method is more difficult to translate languages that lack parallel corpora. Summary of the invention
[0003] The embodiments of the present application provide a text processing method, apparatus, computer device and computer-readable storage medium, which can reduce noise introduced by irrelevant information and improve the processing quality of text information to be processed.
[0004] The technical solution adopted by the present invention to solve the problem is as follows:
[0005] On the one hand, the present application provides a text processing method, comprising:
[0006] Acquire text information to be processed and first associated image information of the text information to be processed;
[0007] Performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information;
[0008] The alignment information is decoded to obtain target text information.
[0009] In some embodiments of the present application, performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information includes:
[0010] Encoding the text information to be processed to obtain first encoded information;
[0011] Extracting features from the first associated image information to obtain visual representation information;
[0012] An alignment operation is performed on the first encoding information and the visual representation information to obtain target alignment information.
[0013] In some embodiments of the present application, feature extraction is performed on the first associated image information to obtain visual representation information, including:
[0014] Segmenting the first associated image information to obtain a plurality of local images;
[0015] Performing projection processing on the plurality of the local images to obtain projection information of the first associated image information;
[0016] Feature extraction is performed on the projection information to obtain visual representation information.
[0017] In some embodiments of the present application, feature extraction is performed on the projection information to obtain visual representation information, including:
[0018] Performing standardization processing on the projection information to obtain standardized projection information;
[0019] Performing feature extraction on the standardized projection information to obtain first feature information;
[0020] Fusing the first feature information with the projection information to obtain first fused information;
[0021] performing standardization processing on the first fusion information to obtain standardized first fusion information;
[0022] Performing feature extraction on the first fusion information after the standardization process to obtain second feature information;
[0023] The second feature information and the first fusion information are fused to obtain visual representation information.
[0024] In some embodiments of the present application, the visual representation information includes global representation information and local representation information, and the aligning operation of the first encoding information and the visual representation information to obtain the target alignment information includes:
[0025] Performing a first alignment operation on the global representation information and the first encoding information to obtain first alignment information;
[0026] Fusing the first alignment information and the local representation information to obtain second fused information;
[0027] A second alignment operation is performed on the second fusion information and the first alignment information to obtain target alignment information.
[0028] In some embodiments of the present application, the text processing method is applied to a text processing model.
[0029] The text processing model is obtained by training a preset network model with a training text set, wherein the preset network model includes a first alignment module, and the first alignment module is used to perform a first alignment operation on the global representation information and the first encoding information;
[0030] Before performing the alignment operation on the text information to be processed and the image information to obtain target alignment information, the method includes:
[0031] Acquire a training text set; the training text set includes a first training text, a second training text, a real text corresponding to the second training text, second associated image information corresponding to the first training text, and third associated image information corresponding to the second training text;
[0032] Inputting the first training text and the second associated image information into a preset network model, and outputting second alignment information through the first alignment module, where the second alignment information is information obtained by performing a first alignment operation on the first training text and the second associated image information;
[0033] Input the second training text, the third associated image information and the real text into a preset network model, and output third alignment information and fourth alignment information through the first alignment module, wherein the third alignment information is information obtained by performing a first alignment operation on the second training text and the third associated image information, and the fourth alignment information is information obtained by performing a first alignment operation on the real text and the third associated image information;
[0034] Determine a first target loss value based on the first alignment information, the second alignment information, and the third alignment information;
[0035] When the first target loss value does not meet the preset first condition, the parameters of the first alignment module are corrected based on the preset first parameter learning rate, and the step of inputting the first training text and the second associated image information into the preset network model is continued until the first target loss value meets the first condition.
[0036] In some embodiments of the present application,
[0037] The preset network model further includes a second alignment module, and the second alignment module is used to perform a second alignment operation on the second fusion information and the first alignment information;
[0038] Before performing the alignment operation on the text information to be processed and the image information to obtain the target alignment information, the method further includes:
[0039] Inputting the first training text and the second associated image information into a preset network model, and outputting fourth alignment information through the second alignment module, where the fourth alignment information is information obtained by performing a second alignment operation on the first training text and the second associated image information;
[0040] Input the second training text, the third associated image information and the real text into a preset network model, and output fifth alignment information and sixth alignment information through the second alignment module, wherein the fifth alignment information is information obtained by performing a second alignment operation on the second training text and the third associated image information, and the sixth alignment information is information obtained by performing a second alignment operation on the real text and the third associated image information;
[0041] Determine a second target loss value based on the first target loss value, the fourth alignment information, the fifth alignment information, and the sixth alignment information;
[0042] When the second target loss value does not meet the preset second condition, the parameters of the second alignment module are corrected based on the preset second parameter learning rate, and the step of inputting the first training text and the second associated image information into the preset network model is continued until the second target loss value meets the second condition.
[0043] In a second aspect, an embodiment of the present invention further provides a text processing device, comprising:
[0044] An information acquisition unit, used for acquiring text information to be processed and first associated image information of the text information to be processed;
[0045] An information alignment unit, used for performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information;
[0046] The text processing unit is used to decode the target alignment information to obtain target text information.
[0047] In a third aspect, the present application further provides a computer device, the computer device comprising:
[0048] one or more processors;
[0049] Memory; and
[0050] One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement any one of the text processing methods in the first aspect.
[0051] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and the computer program is loaded by a processor to execute the steps in any one of the text processing methods in the first aspect.
[0052] The beneficial effects of the present invention are as follows: an alignment operation is performed on the text information to be processed and the first associated image information to obtain target alignment information, and the target alignment information is decoded to obtain target text information. The alignment operation shortens the distance between the corresponding text and image, and based on the image information as a hub, the mapping of low-resource language to target language can be easily achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 is a flowchart of a text processing method provided by an embodiment of the present invention;
[0055] Figure 2 It is a flowchart of a specific embodiment of the text processing method provided by an embodiment of the present invention;
[0056] Figure 3 is a structural diagram of a text processing model provided by an embodiment of the present invention;
[0057] Figure 4 is another structural schematic diagram of a text processing model provided by an embodiment of the present invention;
[0058] Figure 5 is a principle block diagram of a text processing device provided by an embodiment of the present invention;
[0059] Figure 6 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0061] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second", "third", and "fourth" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", "third", and "fourth" may explicitly or implicitly include one or more features. In the description of the present application, "multiple" means two or more, and "several" means one or more, unless otherwise clearly and specifically defined.
[0062] In this application, the word "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any technician in the field to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present application.
[0063] It should be noted that since the method of the embodiment of the present application is executed in a computer device, the processing objects of each computer device exist in the form of data or information. For example, time is actually time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data for processing by the computer device. The details will not be repeated here.
[0064] First, we will briefly explain the professional terms such as parallel corpus, high-resource language and low-resource language. Parallel corpus refers to texts written in different languages that have a "translation relationship" with each other. High-resource language refers to non-scarce language, that is, language with a large population, for which there is a large amount of parallel corpus. The training of machine translation models in high-resource language scenarios has good training samples, and the translation performance of the trained high-resource translation model is also good. Examples of high-resource languages include: English, Chinese, German, etc. Low-resource language refers to scarce language, which is a language with a small population. Its parallel corpus is relatively scarce and the diversity of the corpus is poor, so it cannot be used to train low-resource translation models well. The trained low-resource translation model is prone to overfitting. Examples of low-resource languages include: Turkish, Tujia, etc.
[0065] However, in actual scenarios, due to the uneven development between countries and cultures, most of the world's languages lack parallel corpora, making low-resource languages difficult to translate.
[0066] Based on this, in an embodiment of the present application, the text information to be processed and the first associated image information of the text information to be processed are obtained; the text information to be processed and the first associated image information are aligned to obtain target alignment information; the target alignment information is decoded to obtain target text information. Since the image information and the text information are aligned in the target alignment information, the image information can be used as a hub to easily achieve mapping from a low-resource language to a target language.
[0067] The content of this application is further illustrated below through the description of embodiments in conjunction with the accompanying drawings.
[0068] This embodiment provides a text processing method, such as Figure 1 and Figure 2 As shown in , the method includes:
[0069] Step S10: Acquire the text information to be processed and the first associated image information of the text information to be processed.
[0070] Specifically, the text information to be processed is text that needs to be translated. The text information to be processed can be a sentence or a paragraph input by a user through a computer device (for example, a smart phone), or it can be text data obtained by converting speech to text of collected user conversations through a computer device.
[0071] The first associated image information of the text information to be processed is a number of pictures in random order, and some or all of the pictures are related to the text information to be processed. The first associated image information may be a picture input by the user through a computer device. For example, when the text information to be processed contains words such as "football", "basketball", "sports field", etc., then the pictures related to the text information to be processed may be pictures of scenes such as games or sports. For another example, when the application scenario is video subtitle translation, the text information to be processed is the subtitles that need to be translated, and the first associated image information of the text information to be processed is a picture captured from the video clip corresponding to the subtitles.
[0072] Step S20: performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information.
[0073] The target alignment information is obtained by aligning the encoding information in the text information to be processed and the visual representation information in the first associated image information. The alignment operation shortens the distance between the corresponding text and image. Based on the image information as the hub, the mapping from the low-resource language to the target language can be easily achieved.
[0074] In a specific implementation, Figure 2 As shown, step S20 includes:
[0075] Step S21, encoding the text information to be processed to obtain first encoded information;
[0076] Step S22: extracting features from the first associated image information to obtain visual representation information;
[0077] Step S23: perform an alignment operation on the first encoding information and the visual representation information to obtain target alignment information.
[0078] The first coded information is a coded representation of the text information to be processed obtained by encoding the text information to be processed. The text information to be processed can be represented as X=(x1, x2, ..., x n ), the first coded information can be expressed as w=(w1, w2, ..., w n ), where x n Indicates the nth word in the text information to be processed, w n Represents the text encoding representation of the nth word in the text information to be processed.
[0079] In a specific implementation, the text processing method is applied to a text processing model, such as Figure 3As shown, the text processing model includes a text encoder. Accordingly, the step of encoding the text information to be processed and obtaining the first encoded information of the text information to be processed may specifically include: performing word embedding processing on the text information to be processed to obtain the first word embedding representation of the text information to be processed, adding preset position information to the first word embedding representation to obtain the first word embedding representation with added position information, inputting the first word embedding representation with added position information into the text encoder, and outputting the first encoded information of the text information to be processed through the text encoder.
[0080] Further, continue to refer to Figure 3 As shown, the text encoder includes a first attention layer, a first residual connection layer, a first fully connected layer and a second residual connection layer cascaded in sequence, then the first word embedding representation with added position information is input into the text encoder, and the step of outputting the first encoded information of the text information to be processed through the text encoder can specifically include: inputting the first word embedding representation with added position information into the first attention layer, outputting the first feature through the first attention layer, inputting the first feature and the first word embedding representation with added position information into the first residual connection layer for fusion and normalization processing to obtain the second feature, inputting the second feature into the first fully connected layer for position transformation processing to obtain the second feature after position transformation, and inputting the second feature after position transformation and the second feature into the second residual connection layer for fusion and normalization processing to obtain the first encoded information. Among them, the first fully connected layer can adopt a feedforward neural network, and the first attention layer can adopt a multi-head attention layer.
[0081] In a specific implementation, continue to refer to Figure 2 As shown, step S22 includes:
[0082] Step S221: segment the first associated image information to obtain a plurality of local images.
[0083] Specifically, before the first associated image information is segmented, it is also necessary to encode the first associated image information to obtain image encoding information v0 of the first associated image.
[0084] Furthermore, the first associated image information can be represented as i, and after segmenting the first associated image information i, K two-dimensional local images i are obtained. p , p = 1, 2, ..., K; where, represents the spatial dimension, (H×W) is the resolution of the original image, C is the number of channels, (P×P) is the resolution of the local image, K = HW / P 2 , K represents the number of local images, and K is also the effective input sequence length of the image encoder.
[0085] Step S222: Project the multiple local images to obtain projection information of the first associated image information.
[0086] Specifically, multiple local images are linearly projected to D dimensions respectively, and preset position information is added to obtain the embedded sequence of the local image, and the image coding information v0 is added to the embedded sequence to obtain the embedded sequence z0 after adding the image coding information v0. The linear projection formula is as follows:
[0087]
[0088] Among them, i class Encode information for the image, is the projection information, E pos The projection location information.
[0089] Step S223: extract features from the projection information to obtain visual representation information.
[0090] In a specific implementation, Figure 3 As shown, the text processing model includes an image encoder, which includes an image segmentation submodule, a projection submodule and a feature extraction submodule that are cascaded in sequence. The image is input into the image segmentation submodule, multiple local images are input through the image segmentation submodule, the local images are input into the projection submodule, projection information is input through the projection submodule, the projection information is input into the feature extraction submodule, and visual representation information is output through the feature extraction submodule.
[0091] In a specific implementation, step S223 includes:
[0092] S2231, performing standardization processing on the projection information to obtain standardized projection information;
[0093] S2232, extracting features from the standardized projection information to obtain first feature information;
[0094] S2233, fusing the first feature information with the projection information to obtain first fused information;
[0095] S2234, performing standardization processing on the first fusion information to obtain standardized first fusion information;
[0096] S2235, performing feature extraction on the first fusion information after the standardization process to obtain second feature information;
[0097] S2236. Fuse the second feature information and the first fusion information to obtain visual representation information.
[0098] In a specific implementation, referring to Figure 4 As shown, the feature extraction submodule includes a first standardization module, a first feature extraction module, a fusion module, a second standardization module and a second feature extraction submodule which are cascaded in sequence, the projection information is input into the first standardization module for standardization processing, the standardized projection information is output through the first standardization module, the standardized projection information is input into the first feature extraction module for feature extraction, the first feature information is output through the first feature extraction module, the first feature information and the projection information are input into the first fusion module, the first fusion information is output through the first fusion module, the first fusion information is input into the second standardization module for standardization processing, the first standardized first fusion information is output through the second standardization module, the standardized first fusion information is input into the second feature extraction module, and the visual representation information is output through the second feature extraction module. Among them, the first standardization processing and the second standardization processing can be whitening processing or image normalization processing, the first feature extraction module can adopt a multi-head attention mechanism, the second feature extraction module can adopt a multi-layer perceptron, and the second feature information is extracted through the multi-layer perceptron.
[0099] The visual representation information includes global visual representation information and local visual representation information. The visual representation information is v=(v0, v1, ..., v K ), where v0 is the global visual representation information, v p =(v1, v2, ..., v K ) is the local visual representation information.
[0100] In a specific implementation, continue to refer to Figure 2 As shown, step S23 includes:
[0101] Step S231: perform a first alignment operation on the global visual representation information and the first encoded information to obtain first alignment information.
[0102] Specifically, the first alignment information w=(w1, w2, ..., w n ) to find the mean:
[0103]
[0104] The global visual representation information is defined as:
[0105] v s =v0
[0106] For the global visual representation information v sA first alignment operation is performed on the first encoding information, and the text to be processed is aligned with the first associated image as a whole through the first alignment operation, and the first alignment information w′ is obtained. The first alignment information changes the position information of the text to be processed relative to the first encoding information, and the distance between the text in the second alignment information and its corresponding image is closer.
[0107] Step S232: Fuse the first alignment information and the local representation information to obtain second fused information.
[0108] Step S233: perform a second alignment operation on the second fusion information and the first alignment information to obtain target alignment information.
[0109] Specifically, the first alignment information w′ and the local visual representation information v t Fusion is performed to obtain the second fusion information v t , and then the second fusion information v t Perform a second alignment operation with the first alignment information w′ to obtain the alignment information w out .
[0110] In a specific implementation, Figure 3 As shown, the text processing model includes an alignment module, and the alignment module includes a first alignment module and a second alignment module. The first encoding information and the global visual representation information are input into the first alignment module, and the first alignment information is output through the first alignment module. The first alignment information and the local visual representation information are input into the second alignment module, and the target alignment information is output through the second alignment module.
[0111] In a specific implementation, the second alignment module includes a second attention layer and an information fusion layer that are cascaded in sequence. The first alignment information and the local visual representation information are input into the second attention layer for fusion, the second fusion information is output through the second attention layer, the second fusion information and the first alignment information are input into the information fusion layer, and the target alignment information is output through the information fusion layer.
[0112] Specifically, the second attention layer can use the selective attention mechanism to fuse the second text information with the local visual representation information to obtain the second fused information. The calculation process of the second fused information can be expressed as:
[0113]
[0114] Where w′=(w1′, w2′, ..., w n ′), w′ is the first alignment information, v p =(v1, v2, ..., v K ), v p is the local visual representation information, vt is the second fusion information, W Q , W K and W v is the parameter matrix that needs to be learned during the attention selection operation, w′ is the query for selecting the attention mechanism, and v p The keys and values for selecting the attention mechanism.
[0115] Specifically, the information fusion layer may use a gating mechanism to perform a second alignment operation on the second fusion information and the first alignment information, and the step of aligning the second fusion information and the first alignment information may further include: aligning the second fusion information with the first alignment information to obtain third fusion information, applying the third fusion information to a preset activation function to obtain a gated fusion parameter, and then calculating the target alignment information based on the gated fusion parameter, the first alignment information, and the third fusion information. The activation function may be set as needed. For example, when the activation function is a sigmoid function, the calculation process of the third fusion information may be expressed as follows:
[0116] ε=Sigmoid(W·w′+V·v t )
[0117] Wherein, W and V are learnable parameters. The calculation process of obtaining the target alignment information by calculating the fusion parameters, the first alignment information and the second fusion information can be expressed as:
[0118] w out =(1-ε)·w+ε·v t
[0119] Among them, w out is the target alignment information, and ε is the gated fusion parameter ε∈[0, 1].
[0120] In a specific implementation, the text processing method provided in this embodiment is applied to a text processing model, the text processing model is obtained by training a preset network model through a training text set, the preset network model includes a first alignment module and a decoder, the first alignment module is used to perform a first alignment operation on the global representation information and the first encoding information, and before step S231, it also includes:
[0121] A training text set is obtained, which includes a first training text, a second training text, a real text corresponding to the second text, second associated image information of the first training text, and third associated image information of the second training text. The first training text is a low-resource language text, and the second training text is a high-resource language text. i The description x is used as the first training set, expressed as Among them, Li =(L1, L2, ..., L M ), the first training text L i Including M languages. The third associated image information i of the second training text and the second training text The description x and the second training text The corresponding real text L y The description y is used as the second training set, expressed as
[0122] The first alignment module is trained based on the training text set to obtain the first target loss value, specifically including:
[0123] The first training text and the second associated image information of the first training text are input into a preset network model, and the second alignment information is output through the first alignment module. The second alignment information is information obtained by performing a first alignment operation on the first training text and the second associated image information of the first training text, and the first loss function is calculated based on the second alignment information.
[0124] Specifically, the process of calculating the first loss function is as follows: Since the first training text contains M different languages, it is necessary to divide it into multiple small batches according to the different languages, and then calculate the loss function of each language separately. The first training text is divided into G batches, and the first training text of G batches is encoded to obtain the second encoding information of the first training text: The second associated image information of the first training text is divided into G batches, and the feature extraction of the second associated image information of the first training text of G batches is performed to obtain the global visual encoding information: The first training text will correspond to the second encoding information and the global visual representation information Recorded as a positive example, the remaining G-1 second encoding information and global visual representation information Recorded as negative examples, the first alignment operation makes the representation distance between positive examples closer and the representation distance between negative examples farther. The first loss function of the first alignment module is defined as follows:
[0125]
[0126] in, is the first loss function value, exp is an exponential function with the natural constant e as the base, and s() is the cosine similarity distance, that is, τ is a parameter that determines the difficulty of distinguishing between positive and negative examples. τ The larger the value, the more difficult it is to distinguish between positive and negative examples.
[0127] The second training text, the third associated image information of the second training text, and the real text of the second training text are input into a preset network model, and the third alignment information and the fourth alignment information are output through the first alignment module, and the first probability information is output through the decoder, the third alignment information is information obtained by performing a first alignment operation on the second training text and the third associated image information of the second training text, and the fourth alignment information is information obtained by performing a first alignment operation on the real text of the second training text and the third associated image information.
[0128] The second loss function is calculated according to the third alignment information, the third loss function is calculated according to the fourth alignment information, and the fourth loss function is calculated according to the first probability information. The calculation method of the second loss function, the third loss function and the fourth loss function is similar to the calculation method of the first loss function, and they are all loss functions of the first alignment module, except that the input of the first alignment module is changed.
[0129] The first objective function is calculated according to the first loss function, the second loss function, the third loss function and the fourth loss function. The formula for calculating the first objective function is:
[0130]
[0131] in, represents the first target loss value of the first alignment model, is the first loss function, Indicates that when (i, x) belongs to the first training text set In the middle of the day, expectations; is the second loss function value, Indicates that when (i, x) belongs to the second training text set In the middle of the day, expectations; is the third loss function value, Indicates that when (i, y) belongs to the second training text set In the middle of the day, expectations; is the fourth loss function value, represents the i-th word in the real text corresponding to the second training text, x represents the description of the second training text, y i represents the i-th word in the predicted text corresponding to the second training text, Represents the i-th word y in the predicted text corresponding to the second training text i belong The probability of s is the weight parameter of the first alignment model.
[0132] When the first target loss value does not meet the preset first condition, the parameters of the first alignment module are corrected based on the preset first parameter learning rate, and the step of inputting the first training text and the second associated image information into the preset network model is continued until the first target loss value meets the first preset condition.
[0133] In a specific implementation, the preset network model further includes a second alignment module, and the second alignment module is used to perform a second alignment operation on the second fusion information and the first alignment information. Before step S232, the following steps are also included:
[0134] The first training text and the second associated image information of the first training text are input into a preset network model, and the fifth alignment information is output through the second alignment module. The fifth alignment information is information obtained by performing a second alignment operation on the first training text and the second associated image information of the first training text, and a fifth loss function is calculated based on the fifth alignment information.
[0135] Specifically, the process of calculating the fifth loss function is: extracting features from the second associated image information of the first training text to obtain local visual encoding information. The corresponding second encoding information and global visual representation information in the first training text are extracted. Recorded as a positive example, the remaining second encoding information and global visual representation information Recorded as negative examples, the second alignment operation makes the representation distance between positive examples closer and the representation distance between negative examples farther. The fourth loss function of the second alignment module is defined as follows:
[0136]
[0137] in, is the fifth loss function value, exp is an exponential function with the natural constant e as the base, and s() is the cosine similarity distance, that is, τ is a parameter that measures the difficulty of distinguishing between positive and negative examples. The larger τ is, the more difficult it is to distinguish between positive and negative examples.
[0138] The second training text, the third associated image information of the second training text and the real text of the second training text are input into a preset network model, and the sixth alignment information and the seventh alignment information are output through the second alignment module. The sixth alignment information is information obtained by fusing the second training text and the associated image information of the second training text and performing a second alignment operation, and the seventh alignment information is information obtained by performing a second alignment operation on the real text and the third associated image information.
[0139] The fifth loss function is calculated according to the fifth alignment information, the sixth loss function is calculated according to the sixth alignment information, and the seventh loss function is calculated according to the seventh alignment information. The calculation method of the sixth loss function and the seventh loss function is similar to the calculation method of the fifth loss function, both of which are loss functions of the second alignment module, except that the input of the second alignment module is changed.
[0140] The second objective function is calculated according to the loss value of the first loss function, the fourth loss function, the fifth loss function and the sixth loss function. The formula for calculating the second objective function is:
[0141]
[0142] in, represents the second target loss value, is the fifth loss function value, Indicates that when (i, x) belongs to the first training text set In the middle of the day, expectations; is the sixth loss function value, Indicates that when (i, x) belongs to the second training text set In the middle of the day, expectations; is the seventh loss function value, Indicates that when (i, y) belongs to the second training text set In the middle of the day, The expectation of t is the weight parameter of the second alignment model.
[0143] When the second target loss value does not meet the second preset condition, the second parameter of the second alignment module is corrected based on the preset second parameter learning rate, and the step of inputting the first training text and the second associated image information into the preset network model is continued until the second target loss value meets the second condition.
[0144] In a specific implementation, when the second training text set contains parallel corpora, the parallel corpora in the second training text set are defined as D L ={(x, y)}, then after the correction by the first comparison module and the second comparison module in the above two stages, it is also necessary to use D L Fine-tune the text processing model, and the cross entropy loss function of the fine-tuning process is:
[0145]
[0146] in, represents the loss value of the cross entropy loss function during fine-tuning, represents the i-th word in the real text, x represents the description of the training text, and yi represents the fth word in the predicted text corresponding to the training text, Represents the i-th word y in the predicted text corresponding to the training text i belong The second probability information, Indicates that when x, y belong to the parallel corpus D of the first training text set L In the middle, the formula expectations.
[0147] Step S30: Decode the target alignment information to obtain target text information.
[0148] The target text information is the translated text of the text information to be processed. After obtaining the target alignment information, this embodiment decodes the target alignment information to obtain the translated text of the text information to be processed. Since the target alignment information shortens the distance between the corresponding text and image, based on the image information as the hub, the mapping of low-resource language to the target language can be easily achieved.
[0149] In a specific implementation, continue to refer to Figure 3 As shown, the text processing model also includes a decoder, which decodes the alignment information. The step of obtaining the target text information may specifically include: inputting the alignment information into the decoder, and outputting the target text information through the decoder.
[0150] Continue to refer to Figure 3As shown, the decoder includes a third attention layer, a third residual connection layer, a fourth attention layer, a fourth residual connection layer, a second fully connected layer and a fifth residual connection layer which are cascaded in sequence. Among them, the output item of the third attention layer and the input item of the third attention layer are the input items of the third residual connection layer, the output item of the third residual connection layer and the output item of the alignment module are the input items of the fourth attention layer, the output item of the fourth attention layer and the output item of the third residual connection layer are the input items of the fourth residual connection layer, the output item of the fourth residual connection layer is the input item of the second fully connected layer, and the output item of the second fully connected layer and the output item of the fourth residual connection layer are the input items of the fifth residual connection layer. Among them, the third attention layer can adopt a masked multi-head attention layer, the fourth attention layer can adopt a multi-head attention layer, and the second fully connected layer can adopt a feedforward neural network. The target information is feature extracted through the third attention layer to obtain the third feature information, the third feature information is normalized through the third residual connection layer to obtain the fourth feature information after normalization, the fourth feature information is feature extracted through the fourth attention layer to obtain the fifth feature information, the fourth feature information and the fifth feature information are fused through the fourth residual attention connection layer to obtain the fused feature, the fused feature is linearly activated to obtain the word distribution information, and the target text information is determined based on the word distribution information. The word distribution information is the probability that the distribution of each word in the text is consistent with the distribution of each word in the real text.
[0151] In a specific implementation, the step of outputting the first probability information through the decoder specifically includes: obtaining a training text set, the training text set including a second training text and a real text corresponding to the second training text; inputting the second training text and the real text into a preset network model, and outputting the probability of each word in the predicted text corresponding to the training text belonging to each word in the real text through the preset text processing model; determining the loss value based on the probability and the loss function of the preset network model; when the loss value does not meet the preset conditions, correcting the model parameters of the preset text processing model based on the preset parameter learning rate, and continuing to perform the step of inputting the second training text and the real text into the text processing model until the loss value meets the preset conditions to obtain the text processing model. Among them, the calculation formula of the loss value is: represents the loss value, represents the i-th word in the real text, x represents the second training text, and y i represents the i-th word in the predicted text corresponding to the second training text, Represents the i-th word y in the predicted text corresponding to the second training text i belong probability.
[0152] Continue to refer to Figure 3As shown, during the training of the text processing model, when the real text corresponding to the second training text is input into the text processing model, the real text is first encoded to obtain second encoding information, the position information of the second encoding information is adjusted according to the global visual representation information to obtain second alignment information, and then the local visual representation information and the second alignment information are subjected to a first alignment operation to obtain second fusion information, and the second fusion information and the first alignment information are subjected to a second alignment operation to obtain target alignment information, and then the target alignment information is processed in turn by the second fully connected layer and the fifth residual connection layer, and the processing result is sent to the linear activation layer to obtain the probability that each word in the predicted text corresponding to the second training text belongs to each word in the real text.
[0153] In order to better implement the text processing method in the embodiment of the present application, based on the text processing method, the embodiment of the present application also provides a text processing device, such as Figure 5 As shown, the text processing device includes:
[0154] An information acquisition unit 410 is used to acquire text information to be processed and first associated image information of the text information to be processed;
[0155] An information alignment unit 420 is used to perform an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information;
[0156] The text processing unit 430 is used to decode the target alignment information to obtain target text information.
[0157] In an embodiment of the present application, an alignment operation is performed on text information to be processed and image information associated with the text information to be processed to obtain target alignment information. The alignment operation shortens the distance between the corresponding text and image. Based on the image information as a hub, mapping from a low-resource language to a target language can be easily achieved.
[0158] In some embodiments of the present application, the information alignment unit 420 is specifically used to:
[0159] Encoding the text information to be processed to obtain first encoded information;
[0160] Extracting features from the first associated image information to obtain visual representation information;
[0161] An alignment operation is performed on the first encoding information and the visual representation information to obtain target alignment information.
[0162] In some embodiments of the present application, the information alignment unit 420 is further configured to:
[0163] Segmenting the first associated image information to obtain a plurality of local images;
[0164] Performing projection processing on multiple local images to obtain projection information of associated image information;
[0165] Feature extraction is performed on the projection information to obtain visual representation information.
[0166] In some embodiments of the present application, the information alignment unit 420 is further configured to:
[0167] Performing standardization processing on the first projection information to obtain standardized projection information;
[0168] Performing feature extraction on the standardized projection information to obtain first feature information;
[0169] Fusing the first feature information with the projection information to obtain first fused information;
[0170] Performing standardization processing on the first fusion information to obtain standardized first fusion information;
[0171] Performing feature extraction on the first fusion information after the standardization process to obtain second feature information;
[0172] The second feature information and the first fusion information are fused to obtain visual representation information.
[0173] In some embodiments of the present application, the visual representation information includes global representation information and local representation information, and the information alignment unit 420 is further used to:
[0174] Performing a first alignment operation on the global representation information and the first encoding information to obtain first alignment information;
[0175] Fusing the first alignment information and the local representation information to obtain second fused information;
[0176] A second alignment operation is performed on the second fusion information and the first alignment information to obtain target alignment information.
[0177] The embodiment of the present application further provides a computer device, which integrates any text processing device provided in the embodiment of the present application, and the computer device includes:
[0178] one or more processors;
[0179] Memory; and
[0180] One or more applications, wherein the one or more applications are stored in a memory and are configured to execute, by a processor, the steps of the text processing method in any of the above text processing method embodiments.
[0181] The present application also provides a computer device that integrates any text processing device provided in the present application. Figure 6 As shown, it shows a schematic diagram of the structure of the computer device involved in the embodiment of the present application, specifically:
[0182] The computer device may include one or more processing core processors 501, one or more computer-readable storage media memories 502, a power supply 503, an input unit 504 and other components. Those skilled in the art will appreciate that Figure 6 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:
[0183] The processor 501 is the control center of the computer device. It uses various interfaces and lines to connect various parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 502 and calling data stored in the memory 502, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 501.
[0184] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0185] The computer device also includes a power supply 503 for supplying power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 503 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.
[0186] The computer device may further include an input unit 504, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0187] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 501 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application programs stored in the memory 502, thereby realizing various functions, as follows:
[0188] Acquire the text information to be processed and the first associated image information of the text information to be processed;
[0189] Performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information;
[0190] The target alignment information is decoded to obtain the target text information.
[0191] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0192] To this end, the embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any text processing method provided in the embodiment of the present application. For example, the computer program is loaded by a processor to execute the following steps:
[0193] Acquire the text information to be processed and the first associated image information of the text information to be processed;
[0194] Performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information;
[0195] The target alignment information is decoded to obtain the target text information.
[0196] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.
[0197] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.
[0198] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0199] The above is a detailed introduction to a text processing method, device, computer equipment and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A text processing method, characterized in that: include: Acquire text information to be processed and first associated image information of the text information to be processed; Performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information; The target alignment information is decoded to obtain target text information.
2. The method according to claim 1, characterized in that The aligning operation is performed on the text information to be processed and the first associated image information to obtain target alignment information, including: Encoding the text information to be processed to obtain first encoded information; Extracting features from the first associated image information to obtain visual representation information; An alignment operation is performed on the first encoding information and the visual representation information to obtain target alignment information.
3. The method according to claim 2, characterized in that The extracting features of the first associated image information to obtain visual representation information includes: Segmenting the first associated image information to obtain a plurality of local images; Performing projection processing on the plurality of the local images to obtain projection information of the first associated image information; Feature extraction is performed on the projection information to obtain visual representation information.
4. The method according to claim 3, characterized in that The extracting features of the projection information to obtain visual representation information includes: Performing standardization processing on the projection information to obtain standardized projection information; Performing feature extraction on the standardized projection information to obtain first feature information; Fusing the first feature information with the projection information to obtain first fused information; performing standardization processing on the first fusion information to obtain standardized first fusion information; Performing feature extraction on the first fusion information after the standardization process to obtain second feature information; The second feature information and the first fusion information are fused to obtain visual representation information.
5. The method according to claim 2, characterized in that: The visual representation information includes global representation information and local representation information, and the aligning operation is performed on the first encoding information and the visual representation information to obtain target alignment information, including: Performing a first alignment operation on the global representation information and the first encoding information to obtain first alignment information; Fusing the first alignment information and the local representation information to obtain second fused information; A second alignment operation is performed on the second fusion information and the first alignment information to obtain target alignment information.
6. The method according to claim 1, characterized in that The decoding process of the target alignment information to obtain target text information includes: Extracting features from the target information to obtain third feature information; Performing standardization processing on the third characteristic information to obtain standardized fourth characteristic information; Performing feature extraction on the fourth feature information to obtain fifth feature information; fusing the fourth feature information and the fifth feature information to obtain a fused feature; Perform linear activation processing on the fused features to obtain word distribution information; Based on the word distribution information, target text information is determined.
7. The method according to claim 5, characterized in that The text processing method is applied to a text processing model, the text processing model is obtained by training a preset network model through a training text set, and the preset network model includes a first alignment module and a second alignment module; Before performing the alignment operation on the text information to be processed and the image information to obtain target alignment information, the method includes: Acquire a training text set; the training text set includes a first training text, a second training text, a real text corresponding to the second training text, second associated image information corresponding to the first training text, and third associated image information corresponding to the second training text; Training the first alignment module based on the training text set to obtain a first target loss value; Inputting the first training text and the second associated image information into the preset network model, and outputting fifth alignment information through the second alignment module, where the fifth alignment information is information obtained by performing a second alignment operation on the first training text and the second associated image information; Input the second training text, the third associated image information and the real text into the preset network model, and output sixth alignment information and seventh alignment information through the second alignment module, wherein the sixth alignment information is information obtained by performing a second alignment operation on the second training text and the third associated image information, and the seventh alignment information is information obtained by performing a second alignment operation on the real text and the third associated image information; Determine a second target loss value based on the first target loss value, the fifth alignment information, the sixth alignment information, and the seventh alignment information; When the second target loss value does not meet the preset second condition, the parameters of the second alignment module are corrected based on the preset first parameter learning rate, and the step of inputting the first training text and the second associated image information into the preset network model is continued until the second target loss value meets the second condition.
8. A text processing device, characterized in that: include: An information acquisition unit, used for acquiring text information to be processed and first associated image information of the text information to be processed; An information alignment unit, used for performing an alignment operation on the text information to be processed and the first associated image information to obtain target alignment information; The text processing unit is used to decode the target alignment information to obtain target text information.
9. A computer device, characterized in that: The computer device comprises: one or more processors; Memory; and One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the processor to implement the text processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the text processing method according to any one of claims 1 to 7.