A translation model

By introducing translation models of dynamic convolutional layer, multi-head attention layer and feedforward network layer into the translation model, the rapid accuracy problem of massive image recognition description is solved, and more efficient and accurate image recognition description is achieved.

CN113762408BActive Publication Date: 2025-07-01BEIJING KINGSOFT DIGITAL ENTERTAINMENT CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111087963.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-07-09
Publication Date
2025-07-01
Estimated Expiration
2039-07-09

AI Technical Summary

Technical Problem

When processing massive images, manual identification of descriptions becomes impractical, and a quick and accurate method is needed to achieve image identification descriptions.

Method used

A translation model is provided, including an encoder and a decoder, which includes a dynamic convolutional layer, a multi-head attention layer and a feedforward network layer, through which the image recognition description task is processed.

Benefits of technology

This model can accurately and efficiently perform word processing, and at the same time combines the local feature information of the picture to improve the accuracy of picture recognition, so that the Transformer model can generate more accurate picture descriptions faster in the picture recognition description task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113762408B_ABST
    Figure CN113762408B_ABST
Patent Text Reader

Abstract

The present application provides a translation model and a data processing method. The translation model includes an encoder and a decoder. The decoder includes at least two decoding layers, and at least one of the at least two decoding layers includes a dynamic convolutional layer, a multi-head attention layer, and a feed-forward network layer. Among them, the translation model is used to process picture recognition description tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a translation model and a data processing method. Background Art

[0002] In practical applications, it is often necessary to identify and describe pictures. For example, when classifying pictures, it is necessary to identify the content in the pictures, such as whether it is a scene, an animal, or a person, etc.

[0003] When there are fewer pictures, pictures can be manually identified and described. However, with the development of network technology, the number of pictures has increased sharply. When it is necessary to identify and describe a large number of pictures, the manual processing method becomes too impractical.

[0004] Therefore, how to quickly and accurately identify and describe pictures becomes particularly important. Summary of the Invention

[0005] In view of this, embodiments of this application provide a translation model and a data processing method to solve the technical defects existing in the prior art.

[0006] According to the first aspect of the embodiments of this application, a translation model is provided, which includes an encoder and a decoder. At least two decoding layers are included in the decoder, and at least one of the at least two decoding layers includes a dynamic convolution layer, a multi-head attention layer, and a feed-forward network layer. Among them, the translation model is used to process picture recognition and description tasks.

[0007] Optionally, the dynamic convolution layer, the multi-head attention layer, and the feed-forward network layer are connected in sequence.

[0008] Optionally, the dynamic convolution layer includes a gated linear unit, a dynamic convolution unit, and a lightweight convolution unit.

[0009] Optionally, the gated linear unit is used to receive the matrix to be decoded of the reference picture and obtain a gated linear matrix according to the matrix to be decoded of the reference picture;

[0010] The dynamic convolution unit is used to receive the gated linear matrix and obtain convolution weights according to the gated linear matrix;

[0011] The lightweight convolution unit is used to receive the gated linear matrix and the convolution weights, and obtain a first sub-layer matrix through lightweight convolution operation.

[0012] Optionally, the multi-head attention layer is used to receive the first sub-layer matrix output by the dynamic convolution layer and the picture encoding matrix output by the encoder, and obtain a second sub-layer matrix through multiple self-attention calculations.

[0013] Optionally, the feed-forward network layer is configured to receive the second sub-layer matrix output by the multi-head attention sub-layer, obtain a third sub-layer matrix through feed-forward calculation, and perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix.

[0014] Optionally, the encoder is configured to receive an initial picture matrix to be encoded, and encode the initial picture matrix to be encoded to obtain a picture encoding matrix.

[0015] Optionally, the decoder includes 6 decoding layers.

[0016] Optionally, each decoding layer includes a dynamic convolution layer, a multi-head attention layer, and a feed-forward network layer.

[0017] According to a second aspect of the embodiments of the present application, a data processing method is provided. The method applies the above translation model, and the method includes:

[0018] Based on the received picture to be recognized, an initial picture matrix to be encoded is obtained;

[0019] The initial picture matrix to be encoded is input into the encoder of the translation model for encoding to obtain a picture encoding matrix;

[0020] The picture encoding matrix is input into the decoder of the translation model for decoding to obtain a picture decoding matrix;

[0021] The picture decoding matrix is subjected to normalization processing, and the description information of the picture decoding matrix is output.

[0022] The translation model provided by the present application can effectively combine the local feature information of pictures while accurately and efficiently performing text processing, improve the accuracy of picture recognition, and enable the Transformer model to generate more accurate picture descriptions faster in picture recognition and description tasks. Description of the Drawings

[0023] Figure 1 is a schematic structural diagram of a computing device according to an embodiment of the present application;

[0024] Figure 2 is a schematic flowchart of a data processing method according to an embodiment of the present application;

[0025] Figure 3a is a schematic structural diagram of a dynamic convolution layer according to an embodiment of the present application;

[0026] Figure 3b is a schematic flowchart of a data processing method according to an embodiment of the present application;

[0027] Figures 4a to 4bIt is the architecture diagram of the translation model according to an embodiment of the present application;

[0028] Figure 5 It is the framework schematic diagram of the data processing device according to an embodiment of the present application. Detailed implementation manners

[0029] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0030] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.

[0031] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0032] First, the noun terms related to one or more embodiments of the present invention are explained.

[0033] Transformer: A translation model proposed by Google, which replaces the long short-term memory model with the structure of the self-attention model and has achieved better results in the translation task.

[0034] Self-attention: The attention mechanism is often used in the network structure of the encoder-decoder. It essentially comes from the human visual attention mechanism. When people's vision perceives things, generally not the entire scene is seen, but often a specific part is observed and noted according to the need. The attention mechanism allows the decoder to select the required part from multiple context vectors, thereby representing more information. Taking the decoding layer as an example, for the case where the input vector only comes from the decoding layer itself, it is the self-attention mechanism.

[0035] Multi-head Attention: Also known as Encoder-Decoder Attention. Taking the decoding layer as an example, for the case where the input vectors come from the decoding layer and the encoding layer respectively, it is a multi-head attention mechanism.

[0036] In this application, a data processing method, apparatus, computing device, computer-readable storage medium, and chip are provided, and will be described in detail one by one in the following embodiments.

[0037] Figure 1 FIG. shows a structural block diagram of a computing device 100 according to an embodiment of the present application. The components of the computing device 100 include, but are not limited to, a memory 110 and a processor 120. The processor 120 is connected to the memory 110 through a bus 130, and a database 150 is used to store data.

[0038] The computing device 100 further includes an access device 140, and the access device 140 enables the computing device 100 to communicate via one or more networks 160. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 140 may include one or more of any type of wired or wireless network interfaces (e.g., a Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0039] In an embodiment of the present application, the above components of the computing device 100 and Figure 1 other components not shown in Figure 1 may also be connected to each other, for example, through a bus. It should be understood that

[0040] The structural block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.

[0041] Among them, the processor 120 may execute Figure 2 the steps in the method shown. Figure 2 FIG. shows a schematic flowchart of a data processing method according to an embodiment of the present application. The data processing method of this embodiment is used for a decoder, and the decoder includes at least two decoding layers. For each decoding layer, the method includes the following steps 202 to step 212:

[0042] Step 202: Receive a reference picture to be decoded matrix and a picture encoding matrix.

[0043] Among them, the reference picture to be decoded matrices received for different decoding layers are different. For the first decoding layer, the received reference picture to be decoded matrix is the received initial picture to be decoded matrix; for other decoding layers except the first decoding layer, the received reference picture to be decoded matrix is the picture decoding matrix of the previous decoding layer.

[0044] It should be noted that the initial picture to be decoded matrix is a preset picture decoding matrix.

[0045] Before the first decoding layer of the decoder, it further includes:

[0046] Receive a picture to be recognized;

[0047] Process the picture to be recognized through a pre-trained neural network to obtain a picture feature matrix;

[0048] Perform position encoding on the picture feature matrix to obtain an initial picture to be encoded matrix;

[0049] The encoder receives the initial picture to be encoded matrix and encodes the initial picture to be encoded matrix to obtain a picture encoding matrix.

[0050] Taking the recognition description of a picture as an example, receive a picture to be recognized, and the description information of the picture to be recognized is "a diver observing a sea turtle at the bottom of the sea". Input the picture to be recognized into a pre-trained convolutional application network model to obtain a picture feature matrix; configure an encoding at a corresponding position for each picture feature matrix to obtain an initial picture to be encoded matrix; the encoder receives the initial picture to be encoded matrix and encodes the initial picture to be encoded matrix to obtain a picture encoding matrix, and the decoder receives the picture encoding matrix.

[0051] Step 204: Input the reference picture to be decoded matrix into a dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix.

[0052] The dynamic convolution layer performs convolution operations on the first sub-layer matrix by adopting a parameter sharing mechanism for weights, thereby achieving the purpose of reducing the size of input parameters by the dynamic convolution layer.

[0053] The dynamic convolution calculation process of the reference picture to-be-decoded matrix is shown in the following formula (1);

[0054] DynamicConv(x) = Conv(Linear(x)) (1)

[0055] where x represents the reference picture to-be-decoded matrix, Linear represents a linear mapping, and Conv represents a convolution operation.

[0056] DynamicConv represents the first sub-layer matrix obtained after the dynamic convolution calculation of the reference picture to-be-decoded matrix.

[0057] Figure 3a FIG. is a schematic structural diagram of a dynamic convolution layer, including a gated linear unit, a dynamic convolution unit, and a lightweight convolution unit. Inputting the reference picture to-be-decoded matrix into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix, see Figure 3b , step 204 can be implemented by the following steps 302 to 306:

[0058] Step 302: The gated linear unit receives the reference picture to-be-decoded matrix and obtains a gated linear matrix according to the reference picture to-be-decoded matrix.

[0059] The gated linear unit receives the reference picture to-be-decoded matrix and obtains a gated linear matrix after processing. The gated linear matrix not only effectively reduces gradient dispersion but also retains the non-linear ability.

[0060] Step 304: The dynamic convolution unit receives the gated linear matrix and obtains convolution weights according to the gated linear matrix.

[0061] The dynamic convolution unit receives the gated linear matrix, and the gated linear matrix undergoes dynamic convolution calculation to dynamically generate specific filtering parameters;

[0062] Perform a dot product according to the linear decoding matrix and the filtering parameters, and its output is used as the convolution weights.

[0063] Step 306: The lightweight convolution unit receives the gated linear matrix and the convolution weights, and obtains a first sub-layer matrix through lightweight convolution operation.

[0064] Input the gated linear matrix and the matrix weights into the lightweight convolution unit for lightweight convolution calculation, and perform lightweight convolution operation on the gated linear matrix with the matrix weights to obtain a first sub-layer matrix.

[0065] In an embodiment of the present application, a lightweight convolution operation is performed on the feature matrix of the picture. The weight size is 3×3, and its input channel is 16 and the output channel is 16.

[0066] The number of parameters of the standard convolution operation is 16×16×3×3 = 2304 parameters.

[0067] The lightweight convolution operation realizes spatial convolution by separating the channels into multiple sub-channels and sharing parameters on the sub-channels. The input channel is divided into 4 input sub-channels, and the output channel is divided into 8 output sub-channels. The parameters in the sub-channels are shared. The weights of 4 sizes of 3×3 are traversed through 4 input sub-channels to obtain 4 feature maps, and 8 1×1 are used to traverse these 4 feature maps for fusion. In this process, 4×3×3 + 8×1×1 = 44 parameters are used. Compared with the standard convolution operation, the number of parameters is greatly reduced.

[0068] Step 206: Input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain the second sub-layer matrix.

[0069] The multi-head attention layer performs multiple self-attention calculations on the first sub-layer matrix and the picture encoding matrix to obtain the second sub-layer matrix.

[0070] Step 208: Input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain the third sub-layer matrix.

[0071] The feed-forward network layer can perform feed-forward calculations on the input matrix in parallel and will not further adjust the output result according to the influence of the output result on the input result.

[0072] Step 210: Perform a linear transformation on the third sub-layer matrix to obtain the picture decoding matrix.

[0073] Perform a linear transformation on the obtained third sub-layer matrix. After obtaining the linear matrix, obtain the output picture decoding matrix.

[0074] After obtaining the linear matrix, it is also necessary to perform conventional Residual, Norm, and dropout processing on the linear matrix.

[0075] Residual means that the output of the model is constrained by the residual function to prevent overfitting;

[0076] Norm refers to the normalization operation, which normalizes the output matrix of the model within the range of the normal distribution;

[0077] Dropout means that during the decoding process, the weights of some hidden layer nodes are randomly made not to participate in the work. Those nodes that do not work can be temporarily considered not to be part of the network structure, but their weights need to be retained because they may be required to participate in the work again during the next decoding process.

[0078] Step 212: Output the picture decoding matrix.

[0079] Optionally, use the picture decoding matrix output by the last decoding layer in the decoder as the final picture decoding matrix of the decoder; or perform a fusion calculation based on the picture decoding matrices output by all decoding layers to obtain the final picture decoding matrix of the decoder.

[0080] For a decoder including multiple decoding layers, the final picture decoding matrix of the decoder can be generated by performing a fusion process on the picture decoding matrices of all decoding layers. The fusion method can be to assign weights to the picture decoding matrices of each decoding layer and then sum them to generate the final picture decoding matrix.

[0081] After outputting the picture decoding matrix, it further includes: performing a normalization process on the picture decoding matrix and outputting the description information of the picture decoding matrix.

[0082] Specifically, perform a linear normalization process on the final picture decoding matrix, and the output description information of the picture decoding matrix is "A diver observes a sea turtle at the bottom of the sea", so as to obtain the description information of the picture to be recognized.

[0083] The data processing method provided by this application is used for a decoder. The decoder includes at least two decoding layers; receives a reference picture to-be-decoded matrix and a picture encoding matrix; inputs the reference picture to-be-decoded matrix into a dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; inputs the first sub-layer matrix and the picture encoding matrix into a multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; inputs the second sub-layer matrix into a feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; performs a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; outputs the picture decoding matrix. For each decoding layer, the reference picture to-be-decoded matrix passes through a gated linear unit in the dynamic convolution layer to obtain a gated linear matrix. The dynamic convolution unit dynamically generates the weights in the lightweight convolution operation according to the gated linear matrix. The lightweight convolution unit realizes parameter sharing on multiple sub-channels, reduces the number of parameters, reduces the amount of calculation, reduces the complexity of the algorithm, better focuses on the local feature information of the picture, enables the model to take into account both picture processing and text processing, and can speed up the picture recognition speed while more accurately outputting the description information of the picture.

[0084] For the sake of easy understanding, Figures 4a to 4bThe architecture diagram of the translation model applying the data processing method provided in the embodiments of the present application based on the Transformer model is shown. In the embodiments of the present application, when identifying and describing a picture, the picture to be identified is processed by a pre-trained neural network to obtain a corresponding picture feature matrix. The picture feature matrix is input into the encoder of the Transformer model for encoding processing, and the obtained picture encoding matrix is input into the decoder of the Transformer model. As Figure 4a shown in the Transformer model of

[0085] For each decoding layer, refer to Figure 4b , which includes a dynamic convolution layer, a multi-head attention layer, and a feed-forward network layer. Dynamic convolution, multi-head attention, and feed-forward network are used for calculation respectively to obtain a picture decoding matrix.

[0086] For the first decoding layer: receive the initial picture matrix to be decoded and the picture encoding matrix; input the initial picture matrix to be decoded into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; output the picture decoding matrix.

[0087] For the second decoding layer: receive the picture decoding matrix of the first decoding layer and the picture encoding matrix, input the picture decoding matrix of the first decoding layer into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; output the picture decoding matrix.

[0088] For the third decoding layer: receive the picture decoding matrix of the second decoding layer and the picture encoding matrix, input the picture decoding matrix of the second decoding layer into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; output the picture decoding matrix.

[0089] For the fourth decoding layer: Receive the picture decoding matrix and the picture encoding matrix of the third decoding layer, input the picture decoding matrix of the third decoding layer into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; output the picture decoding matrix.

[0090] For the fifth decoding layer: Receive the picture decoding matrix and the picture encoding matrix of the fourth decoding layer, input the picture decoding matrix of the fourth decoding layer into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; output the picture decoding matrix.

[0091] For the sixth decoding layer: Receive the picture decoding matrix and the picture encoding matrix of the fifth decoding layer, input the picture decoding matrix of the fifth decoding layer into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix; input the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; input the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix; perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix; output the picture decoding matrix.

[0092] Use the picture decoding matrix output by the sixth decoding layer in the decoder as the final picture decoding matrix of the decoder, and perform linear normalization processing on the final picture decoding matrix to obtain the description information of the final picture decoding matrix, thereby obtaining the description information of the picture to be recognized.

[0093] The Transformer model provided by this application can accurately and efficiently perform text processing. Through the dynamic convolution calculation in each decoding layer in the decoder, the operation speed of the model is accelerated, the size of the parameters is reduced, and the convolution operation pays more attention to neighborhood information, enabling the model to more accurately grasp the local feature information of the picture and improving the accuracy of picture recognition. The characteristics of the Transformer model in text processing are jumpiness and neighborhood information. The fusion of the two enables accurate and efficient text processing while effectively integrating the global feature information and local feature information of the picture, enabling the Transformer model to generate more accurate picture descriptions faster in the picture recognition and description task.

[0094] An embodiment of this application further provides a data processing device. SeeFigure 5 , including:

[0095] A first receiving module 502, configured to receive a reference picture to-be-decoded matrix and a picture encoding matrix.

[0096] A dynamic convolution module 504, configured to input the reference picture to-be-decoded matrix into a dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix.

[0097] For the first decoding layer in the decoder, the dynamic convolution module 504 is configured to input the initial picture to-be-decoded matrix as the reference picture to-be-decoded matrix into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix.

[0098] For other decoding layers in the decoder except the first decoding layer; the dynamic convolution module 504 is configured to input the picture decoding matrix of the previous decoding layer as the reference picture to-be-decoded matrix into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix.

[0099] The dynamic convolution module 504 is further configured that the dynamic convolution layer includes a gated linear unit, a dynamic convolution unit, and a lightweight convolution unit; the gated linear unit receives the reference picture to-be-decoded matrix and obtains a gated linear matrix according to the reference picture to-be-decoded matrix; the dynamic convolution unit receives the gated linear matrix and obtains convolution weights according to the gated linear matrix; the lightweight convolution unit receives the gated linear matrix and the convolution weights, and obtains a first sub-layer matrix through lightweight convolution operation.

[0100] A multi-head attention module 506, configured to input the first sub-layer matrix and the picture encoding matrix into a multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix.

[0101] A feed-forward module 508, configured to input the second sub-layer matrix into a feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix.

[0102] A linear module 510, configured to perform a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix.

[0103] An output module 512, configured to output the picture decoding matrix.

[0104] The output module 512 is further configured to use the picture decoding matrix output by the last decoding layer in the decoder as the final picture decoding matrix of the decoder; or perform a fusion calculation based on the picture decoding matrices output by all decoding layers to obtain the final picture decoding matrix of the decoder.

[0105] The normalization module 514 is configured to perform normalization processing on the picture decoding matrix and output the description information of the picture decoding matrix.

[0106] The second receiving module 516 is configured to receive the picture to be recognized.

[0107] The picture processing module 518 is configured to process the picture to be recognized through a pre-trained neural network to obtain a picture feature matrix.

[0108] The position encoding module 520 is configured to perform position encoding on the picture feature matrix to obtain an initial picture matrix to be encoded.

[0109] The encoding module 522 is configured to receive the initial picture matrix to be encoded and encode the initial picture matrix to be encoded to obtain a picture encoding matrix.

[0110] For each decoding layer of the data processing device provided by the present application, the parameter sharing mechanism is adopted for the weights through the dynamic convolution layer in each decoding layer, which can reduce the parameter size. While accelerating the recognition of pictures, the model can generate more accurate picture descriptions.

[0111] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions, and when the instructions are executed by a processor, the steps of the data processing method as described above are implemented.

[0112] The above is a schematic solution of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above data processing method.

[0113] An embodiment of the present application further provides a chip, which stores computer instructions, and when the instructions are executed by the chip, the steps of the data processing method as described above are implemented.

[0114] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0115] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0116] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0117] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this application. The present application selects and specifically describes these embodiments in order to better explain the principle and practical application of the present application, so that those skilled in the art can understand and utilize the present application well. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A translation model, characterized in that, It includes an encoder and a decoder. The decoder includes at least two decoding layers. Each of the at least two decoding layers includes a dynamic convolution layer, a multi-head attention layer, and a feed-forward network layer. Among them, the translation model is used to process picture recognition and description tasks. The dynamic convolution layer includes a gated linear unit, a dynamic convolution unit, and a lightweight convolution unit. Each decoding layer in the translation model receives a reference picture matrix to be decoded and a picture encoding matrix output by the encoder, inputs the reference picture matrix to be decoded into the dynamic convolution layer for dynamic convolution calculation to obtain a first sub-layer matrix, inputs the first sub-layer matrix and the picture encoding matrix into the multi-head attention layer for multi-head attention calculation to obtain a second sub-layer matrix; inputs the second sub-layer matrix into the feed-forward network layer for feed-forward calculation to obtain a third sub-layer matrix, and performs a linear transformation on the third sub-layer matrix to obtain a picture decoding matrix.

2. The translation model according to claim 1, characterized in that, The dynamic convolution layer, the multi-head attention layer, and the feed-forward network layer are connected in sequence.

3. The translation model according to claim 1, wherein The gated linear unit is used to receive the reference picture matrix to be decoded and obtain a gated linear matrix according to the reference picture matrix to be decoded; The dynamic convolution unit is used to receive the gated linear matrix and obtain convolution weights according to the gated linear matrix; The lightweight convolution unit is used to receive the gated linear matrix and the convolution weights, and obtain a first sub-layer matrix through lightweight convolution operation.

4. The translation model according to any one of claims 1 to 3, characterized in that The encoder is used to receive an initial picture matrix to be encoded and encode the initial picture matrix to be encoded to obtain a picture encoding matrix.

5. The translation model according to claim 1, characterized in that The decoder includes 6 decoding layers.

Citation Information

Patent Citations

  • Translation method and device, computing equipment, storage medium and chip

    CN109710953A

  • Dense Video Captioning

    US20190149834A1