Object processing method, device, equipment, storage medium and computer program product
By using row convolution tables for offline and online splitting methods in convolutional neural networks, the problem of inefficient online computing when processing images, text or audio is solved, and a more efficient processing speed is achieved.
Patent Information
- Application Number
- CN202111493729.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-08
AI Technical Summary
Convolutional neural networks have inefficient online computing in image, text or audio processing.
By using the row convolution table to perform offline and online splitting of convolution operations based on pre-trained neural network model, the amount of online operations is reduced. The specific methods include: generating a representation matrix and a row convolution table based on a pre-trained neural network model. When the target convolution kernel slides through the representation matrix, query the row convolution table to obtain the row convolution results, and perform simple addition to obtain the convolution results.
It improves the processing efficiency of input objects, reduces the time spent on online calculation of convolution operations, and improves processing speed.
Smart Images

Figure CN114328802B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer processing technology, and in particular to an object processing method, apparatus, computer equipment, storage medium and computer program product, and also to a text processing method, apparatus, computer equipment, storage medium and computer program product. Background Art
[0002] With the development of artificial intelligence technology, the application of convolutional neural networks has become more and more extensive. Convolutional neural networks can be applied to machine learning models for image processing, text processing, and audio processing. The network structure of a convolutional neural network includes a convolutional layer and a pooling layer, and the convolutional layer is the core of the convolutional neural network.
[0003] Through a large number of convolution kernels in the convolution layer, convolution operations can be performed on the original data of text, image or audio. At present, the convolution operations of the entire convolution layer are all online operations, resulting in low efficiency in processing images, text or audio. Summary of the invention
[0004] Based on this, it is necessary to provide an object processing method, device, computer equipment, storage medium and computer program product that can improve the efficiency of input object processing in response to the above-mentioned technical problems, and also provide a text processing method, device, computer equipment, storage medium and computer program product that can improve the efficiency of input text processing.
[0005] A method for processing an object, the method comprising:
[0006] Determine the representation matrix of the input object based on the pre-trained neural network model;
[0007] Based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding across the representation matrix, determining each representation vector located in the current convolution window from the representation matrix;
[0008] For each representation vector located in the current convolution window, the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index, and a row convolution table generated based on the pre-trained neural network model is queried to obtain a row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel;
[0009] Add the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input object.
[0010] An object processing device, the device comprising:
[0011] A first representation matrix determination module, used to determine a representation matrix of an input object based on a pre-trained neural network model;
[0012] A first convolution kernel sliding module, configured to determine, based on a scale of a target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, each representation vector located within a current convolution window from the representation matrix in a process in which the target convolution kernel slides across the representation matrix;
[0013] A first result query module is used to query the row convolution table generated based on the pre-trained neural network model for each representation vector located in the current convolution window, using the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain a row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel;
[0014] The first result processing module is used to add the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input object.
[0015] In some embodiments, the first representation matrix determination module is used to divide the input object into multiple units; obtain a representation vector set generated based on the pre-trained neural network model; each representation vector in the representation vector set is used to represent each unit that constitutes the object; the representation vectors in the representation vector set are obtained by inputting the units that constitute the object into the pre-trained neural network model; for each unit that constitutes the input object, query the representation vector used to represent the unit from the representation vector set; and determine the representation matrix of the input object based on the representation vector corresponding to each unit.
[0016] In some embodiments, the first representation matrix determination module is further used to determine the distribution of each unit constituting the input object in the input object; and to combine the representation vectors corresponding to each unit according to the distribution to obtain the representation matrix of the input object.
[0017] In some embodiments, the apparatus further comprises: a representation vector representation determination module, configured to determine a representation vector identifier of each representation vector in the representation matrix of the input object;
[0018] The first result query module is also used to query the row convolution table generated based on the pre-trained neural network model for each representation vector located in the current convolution window, using the representation vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain the row convolution result between the representation vector and the row vector at the corresponding row position of the target convolution kernel.
[0019] In some embodiments, the input object is input text, each representation vector in the representation vector set is used to represent each word constituting the text, and the representation vector is a word vector obtained by inputting the word into the pre-trained neural network model for processing the text; or,
[0020] The input object is audio, and each representation vector in the representation vector set is used to represent each segment constituting the audio. The representation vector is an audio vector obtained by inputting the segment into the pre-trained neural network model for processing audio.
[0021] In some embodiments, the device also includes: a row convolution table construction module, which is used to input the various units that make up the sample object into the pre-trained neural network model, obtain the representation vector corresponding to each unit, and generate a unit identifier for each unit or a vector identifier corresponding to each unit; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the representation vector; convolve each representation vector with each convolution kernel row vector to obtain a row convolution result; use the unit identifier or vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index; store each of the third-level index in correspondence with the corresponding row convolution result to generate a row convolution table.
[0022] In some embodiments, the row convolution table construction module is also used to traverse the representation vector and the convolution kernel in the first convolution layer of the pre-trained neural network model; determine the representation vector and the convolution kernel row vector to be convolved; multiply the elements corresponding to the same position in the representation vector and the convolution kernel row vector to obtain the product results of each position; sum the product results of each position to obtain the row convolution result between the representation vector to be convolved and the convolution kernel row vector.
[0023] In some embodiments, the object is text, and a row convolution table construction module is used to input each word constituting the sample text into the pre-trained neural network model, obtain the word vector corresponding to each word, and generate a word identifier or a corresponding word vector identifier for each word; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the word vector; convolve each word vector with each convolution kernel row vector to obtain a row convolution result; use the word identifier or word vector identifier corresponding to the word vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index; store each of the third-level index in correspondence with the corresponding row convolution result to generate a row convolution table.
[0024] In some embodiments, the device also includes: a convolution result determination module, used to determine the current sliding position of the current convolution window in the representation matrix, and use the convolution result corresponding to the current convolution window as the convolution result corresponding to the current sliding position; based on the convolution result corresponding to each sliding position in the process of the target convolution kernel sliding across the representation matrix, generate the convolution result of the representation matrix and the target convolution kernel.
[0025] In some embodiments, the device also includes: a pooling processing module, which is used to perform pooling processing on the convolution result of the representation matrix and the target convolution kernel through a pooling layer in the pre-trained neural network to obtain a pooling result.
[0026] In some embodiments, when the object is an image, the device further includes: an image convolution processing module, which is used to obtain a pixel matrix of an input image; based on the scale of a target convolution kernel to be convolved with the pixel matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding over the pixel matrix, determining each pixel value located in a current convolution window from the pixel matrix; for each pixel value located in the current convolution window, using the pixel value as a first-level index, the convolution kernel identifier of the target convolution kernel as a second-level index, and the position of the pixel value in the current convolution window as a third-level index, querying a product table generated based on the pre-trained neural network model to obtain the pixel value and the target convolution kernel. The product result between the convolution kernel parameters at the position in the convolution kernel; adding the product results between each pixel value in the current convolution window and the convolution kernel parameters at the corresponding position in the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein the product table stores the product result corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model and the position of the convolution kernel parameters, and the product result is obtained by multiplying the pixel value with the convolution kernel parameters at the position in the convolution kernel identified by the convolution kernel identifier; in the process of the target convolution kernel sliding across the pixel matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input image.
[0027] A text processing method, the method comprising:
[0028] Determine a representation matrix of an input text based on a pre-trained neural network model; the representation matrix of the input text is obtained according to a word vector of each word in the input text;
[0029] Based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, determining each word vector located in the current convolution window from the representation matrix in the process of the target convolution kernel sliding across the representation matrix;
[0030] For each word vector located in the current convolution window, the word vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the word vector in the current convolution window is used as the third-level index. The row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel;
[0031] Add the row convolution results between each word vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input text.
[0032] A text processing device, comprising:
[0033] A second representation matrix determination module is used to determine a representation matrix of an input text based on a pre-trained neural network model; the representation matrix of the input text is obtained according to a word vector of each word in the input text;
[0034] A second convolution kernel sliding module is used to determine each word vector located in a current convolution window from the representation matrix in a process in which the target convolution kernel slides across the representation matrix based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model;
[0035] A second result query module is used to query the row convolution table generated based on the pre-trained neural network model for each word vector located in the current convolution window, using the word vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the word vector in the current convolution window as the third-level index, to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel;
[0036] The second result processing module is used to add the row convolution results between each word vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input text.
[0037] A computer device comprises a memory and a processor, wherein the memory stores a computer program and the processor executes the object processing method and / or the text processing method.
[0038] A computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform the above-mentioned object processing method and / or the above-mentioned text processing method.
[0039] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the object processing method and / or the text processing method are implemented.
[0040] In the above-mentioned object processing method, text processing method, object processing device, text processing device and corresponding computer equipment, storage medium, and computer program product, when processing the input object online, in the process of the target convolution kernel sliding over the representation matrix of the input object, for each representation vector located in the current convolution window of the target convolution kernel, the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index. The row convolution table constructed offline in advance is queried to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel. The row convolution result is the result of the offline operation part in the convolution operation. That is to say, the online operation part only needs to obtain the row convolution result by querying the row convolution table, and then simply adding the row convolution result to obtain the convolution result corresponding to the current convolution window. By splitting the convolution operation into two parts, namely, offline operation and online operation, the amount of online operation is reduced, and the time taken for online operation of the convolution operation is reduced, thereby improving the efficiency of online processing of the input object. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 An application environment diagram of an object processing method or a text processing method in an embodiment;
[0042] Figure 2 is a flowchart of an object processing method in one embodiment;
[0043] Figure 3 Schematic diagram of the structure of a convolution kernel in one embodiment;
[0044] Figure 4 is a schematic diagram of a row convolution table based on a three-level index in one embodiment;
[0045] Figure 5 A schematic diagram of summing row convolution results corresponding to a convolution window in one embodiment;
[0046] Figure 6 is a schematic diagram of a row convolution table based on a four-level index in one embodiment;
[0047] Figure 7 A flowchart of a text processing method in one embodiment;
[0048] Figure 8 is a flowchart of a text processing method in another embodiment;
[0049] Fig. 9 is a structural block diagram of an object processing device in some embodiments;
[0050] Fig.10 is a structural block diagram of a text processing device in some embodiments;
[0051] Fig.11 1 is a diagram of the internal structure of a computer device in some embodiments. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] Mentioning "embodiments" in this application means that the specific features, structures or characteristics described in conjunction with the embodiments may be included in at least some embodiments of the present application. The appearance of this phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments. It should be noted that the "first", "second", etc. involved in the introduction of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0054] The object processing method or text processing method provided in this application belongs to the field of artificial intelligence technology. The object processing method provided in this application can be applied to scenarios such as text processing, image processing, and audio processing. The following is an introduction to the terms involved in this application:
[0055] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0056] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0057] Computer Vision (CV) is a science that studies how to make machines "see". To put it more specifically, it refers to the use of cameras and computers to replace human eyes to identify, follow and measure targets, and further perform graphics processing so that the computer processing becomes an image that is more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0058] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS) and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future.
[0059] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0060] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0061] Convolutional Neural Networks (CNN) is a typical network structure of artificial neural networks in artificial intelligence. Convolutional neural networks are mainly used in scenarios such as image, text and audio processing. The network structure of convolutional neural networks is mainly composed of convolutional layers and pooling layers. The convolutional layer is the core computing module of convolutional neural networks.
[0062] The convolution kernel is the core calculation parameter of the convolutional neural network, which has different sizes.
[0063] Convolution operation: The convolution kernel can cover the area of the representation matrix of the input object. The covered area is called the window covered by the convolution kernel. The size of the window is related to the size of the convolution kernel. When the convolution kernel slides over the representation matrix of the input object, the data of the representation matrix of the covered area is multiplied with the parameters of the corresponding position of the convolution kernel to obtain multiple product results, which are then summed to obtain the convolution result corresponding to the window.
[0064] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0065] The object processing method or text processing method provided in this application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through a communication network. Among them, the terminal 102 can be but not limited to a personal computer, a laptop computer, a smart phone, a tablet computer, an Internet of Things device and a portable wearable device. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, a smart car device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented as an independent server or a server cluster consisting of multiple servers.
[0066] In one embodiment, the server 104 may use a data storage system to store a row convolution table, which is mainly used to store a row convolution result corresponding to a representation vector, a convolution kernel identifier of a convolution kernel in a pre-trained neural network model, and a row position of the convolution kernel. The row convolution result is obtained by performing a convolution operation on the representation vector and the convolution kernel row vector corresponding to the convolution kernel identifier and the row position.
[0067] In one embodiment, after the terminal 102 obtains the input object (such as text, image or audio), the input object can be sent to the server 104. The server 104 determines the representation matrix of the input object based on the pre-trained neural network model; based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding over the representation matrix, each representation vector located in the current convolution window is determined from the representation matrix; for each representation vector located in the current convolution window, the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index, and the row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel; the row convolution result between each representation vector in the current convolution window and the row vector of the target convolution kernel corresponding to the corresponding row position is added to obtain the convolution result corresponding to the target convolution kernel in the current convolution window.
[0068] In one embodiment, after the terminal 102 obtains the input text, the input text can be sent to the server 104. The server 104 can determine the representation matrix of the input text based on the pre-trained neural network model; the representation matrix of the input text is obtained according to the word vector of each word in the input text; based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding over the representation matrix, the respective word vectors located in the current convolution window of the target convolution kernel are determined from the representation matrix; for each word vector located in the current convolution window of the target convolution kernel, the word vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the word vector in the current convolution window is used as the third-level index, and the row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the word vector and the row vector of the target convolution kernel corresponding to the row position; the row convolution result between each word vector in the current convolution window and the row vector of the target convolution kernel corresponding to the corresponding row position is added to obtain the convolution result corresponding to the target convolution kernel in the current convolution window.
[0069] In some embodiments, the above-mentioned object processing method and / or text processing method may also be executed by the terminal 102 .
[0070] The object processing method provided in this application can be applied to a computer device (such as a terminal 102 or a server 104). Figure 2 The object processing method provided by the present application is introduced, and the method specifically includes the following steps S201 to S204:
[0071] Step S201, determining a representation matrix of an input object based on a pre-trained neural network model.
[0072] Among them, the pre-trained neural network model is a model obtained by pre-training the model based on training samples. The training samples can be samples based on text, images or audio. The pre-trained neural network model can be used for text recognition, text translation, image recognition, image detection, image segmentation or speech recognition, speech noise reduction, etc.
[0073] In one embodiment, the pre-trained neural network model may include a representation vector generation unit and a convolution feature extraction unit. The representation vector generation unit is mainly used to convert an input object into a representation vector, such as converting an input text into a word vector sequence, or, for example, converting an input image into a vector representing pixels. The convolution feature extraction unit mainly performs a convolution operation on the representation vector generated by the representation vector generation unit to extract features of an object (text, image, or audio, etc.).
[0074] The input object may be an object to be subjected to convolution operation for feature extraction, such as text, image, or audio, etc. The input object may be text, audio, or image, etc. input by a user through the terminal 102 .
[0075] The representation matrix represents the input object in a vectorized mathematical form, and the representation matrix may include multiple representation vectors for representing the input object. Optionally, the representation vectors included in the representation matrix may be generated by a representation vector generation unit of a pre-trained neural network model.
[0076] In one embodiment, after obtaining the input object, the computer device can input the input object into the representation vector generation unit of the pre-trained neural network model, and the representation vector generation unit generates the representation vector of each unit constituting the input object, and arranges and combines the representation vectors of each unit according to the relative position of each unit constituting the input object in the input object to obtain the representation matrix of the input object.
[0077] Taking the input object as text as an example: after the computer device obtains the text input by the user, it can segment the text to obtain each word, and input each word into the representation vector generation unit of the pre-trained neural network model. The representation vector generation unit generates the representation vector of the word, which can be called a word vector. The computer device arranges and combines the word vectors according to the relative position of each word in the text to obtain the representation matrix of the text.
[0078] Take the input object as an example: after the computer device obtains the image input by the user, it can obtain the value of each pixel that makes up the image, and input the value of each pixel into the representation vector generation unit of the pre-trained neural network model for scaling processing at the same proportion to obtain the scaled value of each pixel, and use the scaled value of each pixel as the representation vector of each pixel. According to the relative position of each pixel in the image, the representation vectors of each pixel are arranged and combined to obtain the representation matrix of the image.
[0079] Taking audio as an input object as an example: after the computer device obtains the audio input by the user, it can divide the audio into frames to obtain each voice frame, and input each voice frame into the representation vector generation unit of the pre-trained neural network model, and the representation vector generation unit generates the representation vector of the voice frame. The computer device arranges and combines the representation vectors of the voice frames according to the relative position of each voice frame in the audio to obtain the representation matrix of the audio.
[0080] Step S202, based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding across the representation matrix, each representation vector located in the current convolution window is determined from the representation matrix.
[0081] The target convolution kernel can be any one of the multiple convolution kernels included in the convolution feature extraction unit in the pre-trained neural network model. The target convolution kernel has a certain scale, which is determined by the size of the matrix formed by the convolution kernel parameters. For example, if the matrix formed by the convolution kernel parameters is 2×2, then the scale of the target convolution kernel is 2×2.
[0082] The convolution window corresponds to the scale of the target convolution kernel. For example, if the scale of the target convolution kernel is 2×2, then the convolution window of the target convolution kernel is 2×2. The current convolution window refers to the convolution window that the target convolution kernel currently slides to the position of the representation matrix. When the target convolution kernel slides on the representation matrix of the input object, the representation vectors of the convolution windows at different sliding positions are different.
[0083] When the input object is text, the representation vector located in the convolution window can be a word vector; when the input object is audio, the representation vector located in the convolution window can be a representation vector of a speech frame; when the input object is an image, the representation vector located in the convolution window can be a representation vector of a pixel point of the image.
[0084] In some embodiments, the computer device selects a target convolution kernel from a plurality of convolution kernels included in a convolution feature extraction unit in a pre-trained neural network model. Based on the scale of the target convolution kernel, the computer device determines a representation vector of a convolution window located at a current sliding position when the target convolution kernel slides across a representation matrix of an input object.
[0085] Taking the input object as text as an example, the computer device selects a target convolution kernel from multiple convolution kernels included in the convolution feature extraction unit in the pre-trained neural network model. Based on the scale of the target convolution kernel, the computer device determines the word vector of the convolution window located at the current sliding position when the target convolution kernel slides over the representation matrix composed of word vectors.
[0086] Taking the input object as audio as an example, the computer device selects a target convolution kernel from a plurality of convolution kernels included in the convolution feature extraction unit in the pre-trained neural network model. Based on the scale of the target convolution kernel, the computer device determines the representation vector of the speech frame of the convolution window located at the current sliding position when the target convolution kernel slides over the representation matrix composed of the representation vectors of the speech frame.
[0087] Taking an input object as an image as an example, the computer device selects a target convolution kernel from a plurality of convolution kernels included in a convolution feature extraction unit in a pre-trained neural network model. Based on the scale of the target convolution kernel, the computer device determines the scaled values of the pixels of the convolution window located at the current sliding position when the target convolution kernel slides over a representation matrix composed of scaled values of each pixel.
[0088] Step S203, for each representation vector located in the current convolution window, use the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, query the row convolution table generated based on the pre-trained neural network model, and obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel.
[0089] Among them, the convolution kernel identifier is used to uniquely represent the corresponding convolution kernel. For example, the first convolution kernel can be represented by convolution kernel identifier 1, the second convolution kernel can be represented by convolution kernel identifier 2, the third convolution kernel can be represented by convolution kernel identifier 3, the fourth convolution kernel can be represented by convolution kernel identifier 4, the fifth convolution kernel can be represented by convolution kernel identifier 5, the sixth convolution kernel can be represented by convolution kernel identifier 6, and the seventh convolution kernel can be represented by convolution kernel identifier 7.
[0090] The parameters included in the convolution kernel can be called convolution kernel parameters; the convolution kernel is a matrix composed of convolution kernel parameters. Figure 3 FIG. 1 is a schematic diagram of the structure of a convolution kernel in one embodiment. Figure 3, the column direction of the convolution kernel matrix can be represented by y, and the row direction can be represented by x.
[0091] The row position of the representation vector in the current convolution window is the position of the representation vector along the row direction of the current convolution window. Figure 4 As shown, it is a schematic diagram of a row convolution table based on a three-level index in an embodiment. Taking text as an example, in the convolution window under the current sliding position, the word vectors along the x direction are token2, token3 and token4, respectively. Then the row position of token2 in the current convolution window is the first row, the row position of token3 in the current convolution window is the second row, and the row position of token4 in the current convolution window is the third row. The row position of the representation vector in the current convolution window can be understood as the position of the convolution kernel row vector covering the representation vector in the current convolution window in the convolution kernel. The convolution kernel row vector refers to a vector composed of one or more convolution kernel parameters located in the same row in the convolution kernel.
[0092] The row convolution table is used to store the row convolution results corresponding to the representation vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the row position of the convolution kernel, wherein the row convolution result is obtained by convolving the representation vector with the convolution kernel row vector corresponding to the convolution kernel identifier and the row position. The computer device can construct an index for determining the row convolution result based on the representation vector of the row convolution result, the position of the convolution kernel row vector in the convolution kernel, and the convolution kernel identifier of the convolution kernel to which the convolution kernel row vector belongs; the index can be a three-level index, the first-level index is the representation vector of the row convolution result, the second-level index is the convolution kernel identifier of the convolution kernel to which the convolution kernel row vector of the row convolution result belongs, and the third-level index is the position of the convolution kernel row vector of the row convolution result in the convolution kernel. If the object is text, the constructed three-level index is Figure 4 As shown in [representation vector_index][convolution kernel_index][row position_index]; among them, representation vector_index is used to represent the representation vector, convolution kernel_index is used to represent the convolution kernel, and row position_index is used to represent the position of the convolution kernel row vector in the convolution kernel (it can also be understood as the position of the representation vector in the convolution window).
[0093] The convolution kernel row vector corresponding to the convolution kernel identifier and row position refers to: a vector composed of the convolution kernel parameters at the row position in the convolution kernel corresponding to the convolution kernel identifier. For example, if the convolution kernel identifier is 2 and the row position is the second row, then the convolution kernel row vector corresponding to the convolution kernel identifier 2 and the second row is a vector composed of one or more convolution kernel parameters located in the second row of the convolution kernel of the second convolution kernel.
[0094] Each representation vector belongs to a set of representation vectors generated based on a pre-trained neural network model, that is, each representation vector is generated by a representation vector generation unit of the pre-trained neural network model.
[0095] In the offline stage, each representation vector is convolved with different convolution kernel row vectors in the convolution kernel to obtain each row convolution result.
[0096] like Figure 4 , which is a schematic diagram of a row convolution table based on a three-level index in one embodiment. Figure 4 Take text as an example to introduce the row convolution table.
[0097] Reference Figure 4 , the computer device performs word segmentation on the text, and after obtaining each word, each word is input into the representation vector generation unit of the pre-trained neural network to obtain the representation vector (i.e., word vector) of each word, and each word vector is recorded as token0, token1, token2, ..., tokenN. The computer device can arrange and combine each word vector to obtain the representation matrix of the text, and the representation matrix of the text can be recorded as embedding_dim (dim represents the dimension of the word vector). The convolution kernels included in the convolution feature extraction unit of the pre-trained neural network model are the first convolution kernel, the second convolution kernel, the third convolution kernel, and the fourth convolution kernel. Each word vector is convolved with a different convolution kernel row vector of each convolution kernel to obtain a corresponding row convolution result; the row convolution result corresponds to the representation vector, convolution kernel, and row position of the convolution kernel row vector in the convolution kernel for obtaining the row convolution result. For example, the word vector token0 is convolved with the convolution kernel row vector of the first convolution kernel to obtain a corresponding row convolution result, which corresponds to token0, the convolution kernel identifier 1 of the first convolution kernel, and the first row of the first convolution kernel. The computer device can construct an index to determine the row convolution result as [0][1][0]; wherein the indexes at each level of index [0][1][0] respectively represent that the position of token0, the first convolution kernel, and token0 in the corresponding convolution window is the first row of the first convolution kernel (that is, the convolution kernel row vector convolved with token0 is the convolution kernel row vector of the first row). Similarly, the word vector token0 is convolved with the two convolution kernel row vectors of the second convolution kernel to obtain corresponding row convolution results, where one row convolution result corresponds to token0, the convolution kernel identifier 2 of the second convolution kernel, and the first row of the second convolution kernel, and the other row convolution result corresponds to token0, the convolution kernel identifier 2 of the second convolution kernel, and the second row of the second convolution kernel. The computer device can construct and determine the index of one of the row convolution results as [0][2][0], and construct and determine the index of the other row convolution result as [0][2][1].
[0098] In some embodiments, for each representation vector within the current convolution window of the target convolution kernel, if one of the representation vectors is the word vector token1, and the convolution kernel row vector covering token1 within the current convolution window is the convolution kernel row vector of the third row of the third convolution kernel, then the computer device may use token1 as the first-level index, the convolution kernel identifier 3 of the third convolution kernel as the second-level index, and the row position of token1 in the current convolution window as the third-level index, and the obtained third-level index is exemplarily [1][3][2]. Then, based on the third-level index [1][3][2], the computer device Figure 4 The corresponding row convolution result is determined in the three-level index-row convolution table shown and used as the row convolution result of token1 and the convolution kernel row vector of the third row of the third convolution kernel.
[0099] Step S204, adding the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel, to obtain the convolution result corresponding to the target convolution kernel in the current convolution window.
[0100] In the process of sliding the target convolution kernel over the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input object. For example, if the input object is an image, the representation matrix of the image is processed using the target convolution kernel to extract feature representations such as the edges of the object contours in the image.
[0101] Figure 5 FIG. 1 is a schematic diagram of summing the row convolution results corresponding to the convolution window in one embodiment. Figure 5, the third convolution kernel is used as the target convolution kernel, and the word vectors in the current convolution window of the third convolution kernel are token0, token1, and token2. Among them, the row position of token0 in the current convolution window is the first row, the row position of token1 in the current convolution window is the second row, and the row position of token2 in the current convolution window is the third row. Based on the three-level index introduced in step S203, the computer device can determine the row convolution result between token0 and the first row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by [0][3][0]), determine the row convolution result between token1 and the second row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by index [1][3][1]), determine the row convolution result between token2 and the third row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by index [2][3][2]), add the row convolution results, and use the added result [0][3][0]+[1][3][1]+[2][3][2] as the convolution result corresponding to the current convolution window of the third convolution kernel, which can be used to generate a feature representation of the input text.
[0102] Similarly, if the sliding stride of the third convolution kernel is 1, when the third convolution kernel slides to the next position, the word vectors of the convolution window at the next position are token1, token2 and token3, where the row position of token1 in the convolution window at the next position is the first row, the row position of token2 in the convolution window at the next position is the second row, and the row position of token3 in the convolution window at the next position is the third row. Based on the three-level index introduced in step S203, the computer device can determine the row convolution result between token1 and the first row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by the index [1][3][0]), determine the row convolution result between token2 and the second row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by its index [2][3][1]), determine the row convolution result between token3 and the third row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by its index [3][3][2]), and add the row convolution results, and use the added result [1][3][0]+[2][3][1]+[3][3][2] as the convolution result corresponding to the convolution window at the next position of the third convolution kernel, and the convolution result is used to generate a feature representation of the input text.
[0103] It can be understood that when the convolution feature extraction unit of the pre-trained neural network includes multiple convolution kernels, the computer device can use each convolution kernel to extract features of the input object according to steps S201 to S204 to obtain a feature representation of the input object.
[0104] In the above object processing method, the convolution operation is split into offline and online. In the offline stage, the row convolution table is first used to store the row convolution result corresponding to the representation vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network, and the row position of the convolution kernel. In the online stage, in the process of the target convolution kernel sliding over the representation matrix of the input object, for each representation vector located in the current convolution window of the target convolution kernel, the row convolution table constructed in the offline stage is queried with the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel. The above offline stage performs the multiplication operation of the convolution to obtain the row convolution result. The online stage can obtain the row convolution result by querying the row convolution table, and the online stage only needs to perform a simple addition on the row convolution result to obtain the convolution result corresponding to the current convolution window, thereby improving the speed of online convolution calculation and reducing the time of online convolution calculation.
[0105] In some embodiments, a computer device may determine a representation matrix of an input object through the following steps, specifically including: dividing the input object into multiple units; obtaining a set of representation vectors generated based on a pre-trained neural network model; for each unit constituting the input object, querying a representation vector used to characterize the unit from the representation vector set; and determining the representation matrix of the input object based on the representation vector corresponding to each unit.
[0106] Each representation vector in the representation vector set is used to characterize each unit that constitutes the object; the representation vector in the representation vector set is obtained by inputting the units that constitute the object into a pre-trained neural network model, specifically, it can be obtained by inputting the units that constitute the object into the representation vector generation unit of the pre-trained neural network.
[0107] A unit is a component that can represent an object; by arranging and combining different units, the corresponding object can be obtained. If the object is text, the unit of the text can be a word or a character; if the object is audio, the unit of the audio can be a voice frame; if the object is an image, the unit of the image can be a pixel.
[0108] After obtaining the input object, the computer device divides the input object into multiple units, determines the representation vector of each unit in the representation vector set obtained by the representation vector generation unit based on the pre-trained neural network model, obtains a matrix composed of the representation vectors of each unit according to the relative position of each unit in the input object, and uses the matrix as the representation matrix of the input object.
[0109] In the above embodiment, the representation vector of each unit obtained by segmenting the input object belongs to the representation vector set generated by the pre-trained neural network model. Therefore, it can be ensured that when the target convolution kernel slides on the representation matrix constructed based on the representation vector of each unit, the corresponding row convolution result can be queried through the row convolution table, thereby improving the efficiency of the convolution operation.
[0110] In some embodiments, the computer device may determine the distribution of the various units constituting the input object in the input object; and combine the representation vectors corresponding to each unit according to the distribution to obtain a representation matrix of the input object.
[0111] The distribution of the various units constituting the input object in the input object may be understood as the relative positions of the various units constituting the input object in the input object.
[0112] Take text as an input object as an example: after obtaining the text, the computer device segments the text to obtain the individual words that make up the text. Then, the computer device determines the representation vector corresponding to each word from the representation vector set obtained based on the pre-trained neural network model, obtains the word vector, and combines the word vectors according to the relative position of each word in the text to obtain a matrix, and uses the matrix as the representation matrix of the text.
[0113] In the above embodiment, based on the distribution of the units constituting the input object in the input object, the representation vectors of the units are combined to ensure that the obtained representation matrix can accurately reflect the input object.
[0114] In some embodiments, the computer device may also determine the representation vector identifier of each representation vector in the representation matrix of the input object; for each representation vector located in the current convolution window, the representation vector identifier corresponding to the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index, and the row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel. The representation vector identifier is used to characterize the representation vector.
[0115] Take text as the input object, the target convolution kernel as the third convolution kernel, and the representation vectors located in the current convolution window as token0, token1, and token2 as an example:
[0116] Among them, the row position of token0 in the current convolution window is the first row, that is, the convolution kernel row vector of the first row of the third convolution kernel covers token0; the row position of token1 in the current convolution window is the second row, and the convolution kernel row vector of the second row of the third convolution kernel covers token1; the row position of token2 in the current convolution window is the third row, and the convolution kernel row vector of the third row of the third convolution kernel covers token2. When determining the row convolution result between token0 and the convolution kernel row vector of the first row of the third convolution kernel, the computer device can use token0 as the first-level index, the third convolution kernel as the second-level index, and the first row of the third convolution kernel as the third-level index to obtain the corresponding third-level index [0][3][0]; wherein the indexes of each level of the third-level index [0][3][0] represent token0, the third convolution kernel, and the convolution kernel row vector of the first row of the third convolution kernel, respectively. Based on the three-level index [0][3][0], the computer device can determine the row convolution result between token0 and the convolution kernel row vector of the first row of the third convolution kernel from the row convolution table.
[0117] Similarly, when determining the row convolution result between token1 and the convolution kernel row vector of the second row of the third convolution kernel, the computer device can use token1 as the first-level index, the third convolution kernel as the second-level index, and the second row of the third convolution kernel as the third-level index to obtain the corresponding three-level index [1][3][1]; wherein the three-level indexes of [1][3][1] represent token1, the third convolution kernel, and the convolution kernel row vector of the second row of the third convolution kernel, respectively. The computer device determines the row convolution result between token1 and the convolution kernel row vector of the second row of the third convolution kernel from the row convolution table based on the three-level index [1][3][1].
[0118] When determining the row convolution result between token2 and the convolution kernel row vector of the third row of the third convolution kernel, the computer device can use token2 as the first-level index, the third convolution kernel as the second-level index, and the third row of the third convolution kernel as the third-level index to obtain the corresponding three-level index [2][3][2]; wherein the three-level indexes of [2][3][2] represent token2, the third convolution kernel, and the convolution kernel row vector of the third row of the third convolution kernel, respectively. The computer device determines the row convolution result between token2 and the convolution kernel row vector of the third row of the third convolution kernel from the row convolution table based on the corresponding three-level index.
[0119] In the above embodiment, for each representation vector located in the current convolution window of the target convolution kernel, the identifier of the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index, and the corresponding third-level index is obtained to ensure the accuracy of the row convolution result obtained from the row convolution table query.
[0120] In some embodiments, the input object is input text, each representation vector in the representation vector set is used to represent each word that constitutes the text, and the representation vector is a word vector obtained by inputting the word into a pre-trained neural network model for processing text.
[0121] The representation vector of a word may be obtained by inputting the word into a representation vector generation unit of a pre-trained neural network model and outputting the representation vector generation unit.
[0122] In some embodiments, the input object is audio, each representation vector in the representation vector set is used to represent each segment that constitutes the audio, and the representation vector is an audio vector obtained by inputting the segment into a pre-trained neural network model for processing audio.
[0123] Among them, the various segments that make up the audio can be called speech frames, and the representation vector of the speech frame can be the speech frame input into the representation vector generation unit of the pre-trained neural network model and output by the representation vector generation unit.
[0124] In the above embodiment, the input object may be text or audio. When performing convolution operation on the text or audio, the corresponding row convolution result may be queried through the row convolution table, thereby improving the convolution processing efficiency of text recognition and speech recognition.
[0125] In some embodiments, when constructing a row convolution table, the computer device can input the various units that make up the sample object into a pre-trained neural network model to obtain a representation vector corresponding to each unit, and generate a unit identifier for each unit or a vector identifier corresponding to each unit; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the representation vector; perform convolution operations on each representation vector with each convolution kernel row vector to obtain a row convolution result; use the unit identifier or vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index; store each third-level index in correspondence with the corresponding row convolution result to generate a row convolution table.
[0126] The sample object may be an object used in the process of constructing a pre-trained neural network model; each unit of the sample object is input into the representation vector generation unit of the pre-trained neural network model, and the representation vector generation unit generates the representation vector of each unit. Each unit has a corresponding unit identifier, and the unit identifier is used to uniquely represent the unit. The representation vector of each unit also has a corresponding vector identifier, and the vector identifier is used to uniquely represent the representation vector of the unit.
[0127] In the pre-trained neural network model, the representation vector input to the first convolution layer is output by the representation vector generation unit. The representation vector input to the first convolution layer is relatively fixed. Therefore, the convolution kernel and each representation vector of the first convolution layer can be pre-calculated. The data to be convolved in the deeper convolution layer is not only related to the convolution kernel of the previous layer, but also to the representation vector input to the first convolution layer. There are many influencing factors. At this time, the convolution kernel of the deeper convolution layer and the data to be convolved in the deeper convolution layer do not need to be pre-calculated.
[0128] Assume that the convolution kernels in the first convolution layer are the first convolution kernel, the second convolution kernel, the third convolution kernel, and the fourth convolution kernel, such as Figure 4 As shown, the first convolution kernel includes one convolution kernel row vector, the second convolution kernel includes two convolution kernel row vectors, the third convolution kernel includes three convolution kernel row vectors, and the fourth convolution kernel includes four convolution kernel row vectors.
[0129] When the object is text or audio, the dimension of the representation vector is consistent with the dimension of the convolution kernel row vector of each convolution kernel of the first convolution layer. For example, if the dimension of the convolution kernel row vector is 8 dimensions, the dimension of the convolution kernel row vector is also 8 dimensions.
[0130] Take the text as an example: if the word vector of the text includes token0, token1 and token2, then the computer device can perform convolution operations on token0 with the convolution kernel row vector included in the first convolution kernel, the convolution kernel row vector included in the second convolution kernel, the convolution kernel row vector included in the third convolution kernel, and the convolution kernel row vector included in the fourth convolution kernel, respectively, to obtain the row convolution results of token0 corresponding to different convolution kernel row vectors. Similarly, the computer device can perform convolution operations on token1 with the convolution kernel row vector included in the first convolution kernel, the convolution kernel row vector included in the second convolution kernel, the convolution kernel row vector included in the third convolution kernel, and the convolution kernel row vector included in the fourth convolution kernel, respectively, to obtain the row convolution results of token1 corresponding to different convolution kernel row vectors. The computer device can perform convolution operations on token2 with the convolution kernel row vector included in the first convolution kernel, the convolution kernel row vector included in the second convolution kernel, the convolution kernel row vector included in the third convolution kernel, and the convolution kernel row vector included in the fourth convolution kernel, respectively, to obtain row convolution results of token2 corresponding to different convolution kernel row vectors.
[0131] When a computer device constructs an index for a row of convolution results, it can use the vector identifier or unit identifier of the representation vector of the row of convolution results as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index, and store the third-level index in correspondence with the row of convolution results to generate a row convolution table.
[0132] For example, when the computer device constructs the index of the row convolution result between token0 and the convolution kernel row vector included in the first convolution kernel, the vector identifier of token0 or the unit identifier of the unit corresponding to token0 is used as the first-level index, the convolution kernel identifier 1 of the first convolution kernel is used as the second-level index, and the row position of the convolution kernel row vector in the first convolution kernel is used as the third-level index to generate a third-level index [0][1][0]; the computer device stores the third-level index [0][1][0] in correspondence with the row convolution result between token0 and the convolution kernel row vector included in the first convolution kernel to obtain a row convolution table.
[0133] For each unit constituting the sample object, the computer device inputs each unit into the representation vector generation unit of the pre-trained neural network model, and the representation vector generation unit outputs the representation vector of each unit. The computer device obtains the unit identification of each unit or the vector identification of the representation vector of the unit. For each convolution kernel in the first convolution layer of the pre-trained neural network, the convolution kernel identification of each convolution kernel and the convolution kernel row vector of each convolution kernel are determined; the representation vector of each unit is convolved with each convolution kernel row vector respectively, and the row convolution result of the representation vector of each unit corresponding to different convolution kernel row vectors is obtained. When the computer device constructs the index corresponding to any row convolution result, the vector identification of the representation vector of any row convolution result or the unit identification of the unit corresponding to the representation vector is used as the first-level index, the convolution kernel identification of the convolution kernel where the convolution kernel row vector of any row convolution result is located is used as the second-level index, and the row position of the convolution kernel row vector of any row convolution result in the convolution kernel is used as the third-level index, and the third-level index corresponding to any row convolution result is obtained. The computer device stores the three-level index and the corresponding row convolution result to obtain a row convolution table of the three-level index.
[0134] In the above embodiment, the representation vector of each unit of the sample object is convolved with the convolution kernel row vector of each convolution kernel of the first convolution layer of the pre-trained neural network model to ensure the comprehensiveness of the obtained row convolution result. In addition, the unit representation or vector identifier corresponding to the representation vector is used as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located is used as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel is used as the third-level index to ensure the accuracy of the row convolution result obtained by online query.
[0135] In some embodiments, when the computer device performs convolution operations on each representation vector and each convolution kernel row vector to obtain a row convolution result, it can traverse the representation vector and the convolution kernel in the first convolution layer of the pre-trained neural network model; determine the representation vector and the convolution kernel row vector to be convolved; multiply the elements corresponding to the same position in the representation vector and the convolution kernel row vector to obtain the product result at each position; sum the product results at each position to obtain the row convolution result between the representation vector and the convolution kernel row vector to be convolved.
[0136] Among them, the representation vector may include multiple representation data arranged in order, and the convolution kernel row vector includes multiple convolution kernel parameters arranged in order.
[0137] The elements corresponding to the same position in the representation vector and the convolution kernel row vector may refer to: the representation data covered by the convolution kernel parameters. For example, the ordered representation data of the word vector token0 is [a1 a2 a3 a4 a5], and the ordered arrangement of the convolution kernel parameters of the convolution kernel row vector is [b1 b2 b3 b4 b5], where b1 covers a1, b2 covers a2, b3 covers a3, b4 covers a4, and b5 covers a5. Therefore, a1 and b1 are the elements corresponding to the same position in the word vector token0 and the convolution kernel vector, a2 and b2 are the elements corresponding to the same position in the word vector token0 and the convolution kernel vector, a3 and b3 are the elements corresponding to the same position in the word vector token0 and the convolution kernel vector, a4 and b4 are the elements corresponding to the same position in the word vector token0 and the convolution kernel vector, and a5 and b5 are the elements corresponding to the same position in the word vector token0 and the convolution kernel vector.
[0138] The multiplication of the vector and the elements in the same position in the convolution kernel row vector can be called bitwise product.
[0139] After the computer device obtains the convolution kernel in the first convolution layer of the pre-trained neural network model, the representation vector of each unit is convolved with the convolution kernel row vector of the convolution kernel. When the representation vector of the unit is convolved with any convolution kernel row vector, the representation vector of the unit is bitwise multiplied with the row vector of any convolution kernel to obtain the product result of each position, and the product results of each position are summed, and the result obtained by the summation is the row convolution result between the representation vector of the unit and the row vector of any convolution kernel.
[0140] In the above embodiment, the representation vector to be convolution operation is bitwise multiplied and summed with the convolution kernel row vector to ensure the accuracy of the obtained row convolution result.
[0141] In some embodiments, the object is text, and when the computer device constructs a row convolution table, the various words that make up the sample text can be input into a pre-trained neural network model to obtain the word vector corresponding to each word, and generate a word identifier or a corresponding word vector identifier for each word; for each convolution kernel in the first convolution layer of the pre-trained neural network model, the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel are determined; the dimension of the convolution kernel row vector is the same as the dimension of the word vector; each word vector is convolved with each convolution kernel row vector to obtain a row convolution result; a third-level index is generated using the word identifier or word vector identifier corresponding to the word vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index; each third-level index is stored in correspondence with the corresponding row convolution result to generate a row convolution table.
[0142] The sample text may be the text used in constructing the pre-trained neural network model. After the computer device performs word segmentation on the sample text, each word of the sample text is obtained, and each word is input into the pre-trained neural network model. The representation vector generation unit of the pre-trained neural network model generates a representation vector for each word (the representation vector of each word may be referred to as a word vector); the computer device may generate a corresponding word identifier or word vector identifier for each word vector. The first convolution layer of the convolution feature extraction unit of the pre-trained neural network model includes a plurality of convolution kernels (for convenience of explanation, let Figure 4 The first convolution kernel, the second convolution kernel, the third convolution kernel and the fourth convolution kernel shown are convolution kernels of the first convolution layer), and the dimension of the convolution kernel row vector of each convolution kernel is consistent with the dimension of the above-mentioned word vector. The computer device can determine the convolution kernel row vectors of the first convolution kernel, the second convolution kernel, the third convolution kernel and the fourth convolution kernel, and perform convolution operations on each word vector with each convolution kernel row vector to obtain corresponding row convolution results.
[0143] For any row convolution result, the computer device can use the word vector identifier or the corresponding word identifier of the word vector obtained for any row convolution as the first-level index (which can be recorded as the representation vector_index), use the convolution kernel identifier of the convolution kernel where the convolution kernel row vector of any row convolution result is located as the second-level index (which can be recorded as the convolution kernel_index), and use the row position of the convolution kernel row vector of any row convolution result in the convolution kernel where it is located as the third-level index (which can be recorded as the row position_index), and obtain the third-level index corresponding to the any row convolution result, and the third-level index includes [representation vector_index][convolution kernel_index][row position_index]. The computer device stores the any row convolution result and the three-level index corresponding to the any row convolution result in the row convolution table to obtain the row convolution table applied to text convolution processing.
[0144] In the above embodiment, the word vector of each sample text is convolved with the convolution kernel row vector of the convolution kernel of the first convolution layer of the pre-trained neural network model in advance, and the obtained row convolution result and the corresponding three-level index are stored to ensure that when the text is subsequently subjected to online convolution processing, the accurate row convolution result can be queried from the row convolution table, thereby improving the online processing efficiency.
[0145] In some embodiments, the computer device can also determine the current sliding position of the current convolution window in the representation matrix, and use the convolution result corresponding to the current convolution window as the convolution result corresponding to the current sliding position; based on the convolution result corresponding to each sliding position in the process of the target convolution kernel sliding across the representation matrix, generate the convolution result of the representation matrix and the target convolution kernel.
[0146] The sliding position can be understood as the position of the convolution window in the representation matrix. When the convolution kernel slides on the representation matrix, the convolution window is at different positions in the representation matrix. Therefore, there are multiple sliding positions.
[0147] The computer device can determine the current sliding position of the current convolution window in the representation matrix, and use the convolution result corresponding to the current convolution window as the convolution result of the current sliding position; similarly, when the convolution window slides to the next position, the computer device can determine the next sliding position of the convolution window in the representation matrix, and use the convolution result corresponding to the convolution window at the next sliding position as the convolution result of the next sliding position.
[0148] After obtaining the convolution results corresponding to each sliding position, the computer device arranges the convolution results of each sliding position according to the distribution of each sliding position, and uses the arranged results as the convolution results of the representation matrix and the target convolution kernel.
[0149] In the above embodiment, the convolution result corresponding to the convolution window at each sliding position is used as the convolution result corresponding to each sliding position, thereby ensuring the accuracy of the convolution result of the obtained representation matrix and the target convolution kernel.
[0150] In some embodiments, the computer device may also perform pooling processing on the convolution result of the representation matrix and the target convolution kernel through a pooling layer in a pre-trained neural network to obtain a pooling result.
[0151] The pooling process may be to perform maximum pooling or average pooling using the convolution kernel of the pooling layer. Maximum pooling may be: the convolution result representing the matrix and the target convolution kernel includes the convolution results at multiple sliding positions, and when the convolution kernel of the pooling layer slides on the convolution result representing the matrix and the target convolution kernel, the largest convolution result is selected from the convolution results at multiple sliding positions within the current convolution window of the convolution kernel of the pooling layer as the result corresponding to the current convolution window of the convolution kernel of the pooling layer, to obtain the corresponding pooling result.
[0152] Average pooling can be: the convolution result representing the matrix and the target convolution kernel includes the convolution results at multiple sliding positions. When the convolution kernel of the pooling layer slides on the convolution result representing the matrix and the target convolution kernel, the convolution results at multiple sliding positions within the current convolution window of the convolution kernel of the pooling layer are averaged, and the result of the average processing is used as the result corresponding to the current convolution window of the convolution kernel of the pooling layer to obtain the corresponding pooling result.
[0153] After obtaining the convolution result of the representation matrix and the target convolution kernel, the computer device uses the pooling layer of the pre-trained neural network to perform average pooling or maximum pooling on the convolution result of the representation matrix and the target convolution kernel to obtain the corresponding pooling result.
[0154] In the above embodiment, the pooling layer of the pre-trained neural network is used to perform pooling processing on the convolution result of the representation matrix and the target convolution kernel, thereby reducing the amount of data to be processed and improving processing efficiency.
[0155] In some embodiments, when the object is an image, the computer device can obtain a pixel matrix of the input image; based on the scale of the target convolution kernel to be convolved with the pixel matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding over the pixel matrix, determine each pixel value located in the current convolution window from the pixel matrix; for each pixel value located in the current convolution window, use the pixel value as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the position of the pixel value in the current convolution window as the third-level index, query the product table generated based on the pre-trained neural network model, and obtain the product result between the pixel value and the convolution kernel parameters at the position in the target convolution kernel; add the product results between each pixel value in the current convolution window and the convolution kernel parameters at the corresponding position in the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window.
[0156] Among them, the product table stores the product results corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the position of the convolution kernel parameters. The product result is obtained by multiplying the pixel value with the convolution kernel parameter at the position in the convolution kernel identified by the convolution kernel identifier.
[0157] In the product result corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the position of the convolution kernel parameter, the position of the convolution kernel parameter can be understood as the position of the convolution kernel parameter in the convolution kernel that obtains the product result. Since the convolution kernel is a matrix composed of the convolution kernel parameters, the position of the convolution kernel parameter in the convolution kernel can include the row position of the convolution kernel parameter in the convolution kernel and the column position of the convolution kernel parameter in the convolution kernel; wherein the row position refers to the position of the convolution kernel parameter along the convolution kernel. Figure 3 The x-direction position is shown, and the column position refers to the convolution kernel parameters along Figure 3 The position in the y direction is shown. Correspondingly, the third-level index is subdivided into two sub-level indexes, namely the index of the position along the x direction and the index of the position along the y direction; at this time, in the index corresponding to the product result, the first-level index of the index is recorded as [pixel value_index], the second-level index is recorded as [convolution kernel_index], and the two sub-level indexes of the third-level index are recorded as [row position_index] and [column position_index], respectively. Therefore, the index including [pixel value_index], [convolution kernel_index] and the two sub-level indexes of the third level can be called a four-level index, and the corresponding product table can be called a four-level index-product table.
[0158] After obtaining the pixel matrix of the image, the computer device multiplies each pixel value of the pixel matrix with each convolution kernel parameter of the convolution kernel, and stores each product result and index in correspondence to obtain a product table. Figure 6The above process is described using the pixel value pixel_00 as an example:
[0159] After the computer device obtains the fifth convolution kernel, the sixth convolution kernel, and the seventh convolution kernel, each convolution kernel slides on the pixel value pixel_00 so that each convolution kernel parameter of each convolution kernel is multiplied with the pixel value pixel_00 to obtain the corresponding convolution result. Among them, if the convolution kernel parameter of the first row and first column of the fifth convolution kernel is multiplied with the pixel value pixel_00, then the position of the pixel value pixel_00 on the fifth convolution kernel is the first row and first column. Let the first row be 0 and the first column be 0, then the index corresponding to the product result between the convolution kernel parameter of the first row and first column of the fifth convolution kernel and the pixel value pixel_00 is
[00] [5][0][0], where
[00] represents the pixel value pixel_00, [5] represents the fifth convolution kernel, and [0][0] represents the first row and first column.
[0160] The computer device can store the product results of the pixel value pixel_00 and the convolution kernel parameters of other convolution kernels, and store the product results of other pixel values of the image pixel matrix and the convolution kernel parameters of each convolution kernel in the above manner.
[0161] After obtaining the pixel matrix of the input image, the computer device can input the pixel matrix of the input image into the representation vector generation unit of the pre-trained neural network to perform proportional scaling processing on each pixel value of the pixel matrix.
[0162] Next, the computer device selects a target convolution kernel from the convolution kernel of the convolution feature extraction unit of the pre-trained neural network. When the target convolution kernel slides on the pixel matrix of the input image, each pixel value located in the current convolution window can be determined. For each pixel value located in the current convolution window, the position of each pixel value in the current convolution window is the position of the convolution kernel parameter in the convolution kernel that covers the pixel value and is to be multiplied with the pixel value. Next, the pixel value is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, the row position of the pixel value in the current convolution window and the column position of the pixel value in the current convolution window are used as the sub-level index of the third-level index, and a four-level index with a structure of [pixel value_index][convolution kernel_index][row position_index][column position_index] is constructed. The computer device queries the four-level index-product table according to the constructed index, obtains the product result corresponding to the constructed index, and uses the product result as the product result between the pixel value and the convolution kernel parameter at the position in the target convolution kernel.
[0163] Next, the computer device can obtain the product results between each pixel value within the convolution window and the convolution kernel parameters covering the pixel values, add the product results, and use the added result as the convolution result corresponding to the current convolution window.
[0164] In the process of sliding the target convolution kernel over the pixel matrix of the input image, the convolution results corresponding to each convolution window refer to the convolution results corresponding to the convolution windows at each sliding position. The convolution results corresponding to each convolution window are used to generate the feature representation of the input image. According to the characteristics of the convolution kernel, the feature representation of the input image generated based on the convolution results corresponding to each convolution window can be the edge features of the object in the image and the features after smoothing the noise in the image.
[0165] In the above embodiment, a product table is first used to store the product results corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the position of the convolution kernel parameters. Then, when the pixel matrix of the input image is obtained, the pixel value located in the convolution window is the first-level index as the target convolution kernel slides across the pixel matrix, the convolution kernel identifier of the target convolution kernel is the second-level index, and the position of the pixel value in the current convolution window is the third-level index. The corresponding product result is queried in the product table without performing a product operation, thereby improving the efficiency of convolution processing on the image.
[0166] The present application provides an object processing method, which can be applied to a computer device such as a terminal 102 or a server 104, and the method includes the following steps:
[0167] Step S601, dividing the input object into multiple units;
[0168] Step S602, obtaining a set of representation vectors generated based on a pre-trained neural network model;
[0169] Step S603, for each unit constituting the input object, querying a representation vector for representing the unit from the representation vector set;
[0170] Step S604, determining the distribution of each unit constituting the input object in the input object;
[0171] Step S605, combining the representation vectors corresponding to each unit according to distribution to obtain a representation matrix of the input object;
[0172] Step S606, based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding across the representation matrix, determining each representation vector located in the current convolution window from the representation matrix;
[0173] Step S607, for each representation vector located in the current convolution window, the representation vector identifier corresponding to the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index, and the row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel;
[0174] Step S608, adding the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel, to obtain the convolution result corresponding to the target convolution kernel in the current convolution window;
[0175] Step S609, determining the current sliding position of the current convolution window in the representation matrix, and taking the convolution result corresponding to the current convolution window as the convolution result corresponding to the current sliding position;
[0176] Step S610, generating a convolution result of the representation matrix and the target convolution kernel according to the convolution result corresponding to each sliding position in the process of the target convolution kernel sliding over the representation matrix;
[0177] Step S611, performing pooling processing on the convolution result of the representation matrix and the target convolution kernel through the pooling layer in the pre-trained neural network to obtain a pooling result;
[0178] In the above object processing method, the convolution operation is split into offline and online. In the offline stage, the row convolution table is first used to store the row convolution results corresponding to the word vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network, and the row position of the convolution kernel. In the online stage, in the process of the target convolution kernel sliding over the representation matrix of the input text, for each word vector located in the current convolution window of the target convolution kernel, the word vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the word vector in the current convolution window is used as the third-level index. The row convolution table constructed in the offline stage is queried to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel. The above offline stage performs the product operation of the convolution to obtain the row convolution result. The row convolution result can be obtained by querying the row convolution table in the online stage. In the online stage, only a simple addition of the row convolution result is required to obtain the convolution result corresponding to the current convolution window, which improves the speed of online convolution calculation and reduces the time of online convolution calculation.
[0179] The present application provides a text processing method that can be applied to a computer device such as a terminal 102 or a server 104, such as Figure 7 As shown, the method includes:
[0180] Step S701, determining a representation matrix of the input text based on the pre-trained neural network model.
[0181] Among them, the representation matrix of the input text is obtained based on the word vector of each word in the input text.
[0182] The computer device can segment the input text to obtain the individual words that make up the input text, and input the individual words that make up the input text into the pre-trained neural network model. The representation vector generation unit of the pre-trained neural network model generates the representation vector of each word, and the representation vector of each word can be called a word vector. Then, the computer device arranges the word vectors according to the distribution of each word in the input text to obtain a representation matrix of the input text. Among them, each word vector is arranged along the row direction of the matrix, and a word vector is placed in each row of the representation matrix of the input text; the words corresponding to the two word vectors in adjacent rows are adjacent when the input text is.
[0183] Step S702, based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding across the representation matrix, determine each word vector located in the current convolution window from the representation matrix.
[0184] The target convolution kernel is the convolution kernel included in the convolution feature extraction unit in the pre-trained neural network model. The convolution window corresponds to the scale of the target convolution kernel. For example, if the scale of the target convolution kernel is 2×2, then the convolution window of the target convolution kernel is 2×2.
[0185] The computer device can determine a target convolution kernel in the convolution kernel of the convolution feature extraction unit of the pre-trained neural network model; in the process of the target convolution kernel sliding across the representation matrix of the input text, determine each word vector located in the current convolution window in the representation matrix.
[0186] For other related introductions, please refer to the contents of other embodiments.
[0187] Step S703, for each word vector located in the current convolution window, use the word vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the word vector in the current convolution window as the third-level index, query the row convolution table generated based on the pre-trained neural network model, and obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel.
[0188] Among them, the convolution kernel identifier is mainly used to uniquely characterize the corresponding convolution kernel, for example Figure 4 The first convolution kernel shown can be represented by convolution kernel identification 1, the second convolution kernel can be represented by convolution kernel identification 2, the third convolution kernel can be represented by convolution kernel identification 3, and the fourth convolution kernel can be represented by convolution kernel identification 4.
[0189] The parameters included in the convolution kernel can be called convolution kernel parameters; the convolution kernel is a matrix composed of convolution kernel parameters (such as Figure 3 As shown in Figure 2), the column direction in the matrix of the convolution kernel can be represented by y, and the row direction can be represented by x.
[0190] The row position of the word vector in the current convolution window is the position of the word vector along the row direction of the convolution window at the current sliding position. For example, in the convolution window at the current sliding position, the word vectors along the x direction are token2, token3, and token4, then the row position of token2 in the current convolution window is the first row, the row position of token3 in the current convolution window is the second row, and the row position of token4 in the current convolution window is the third row.
[0191] The row position of the word vector in the current convolution window can be understood as the position of the convolution kernel row vector covering the word vector in the current convolution window. The convolution kernel row vector refers to the vector composed of one or more convolution kernel parameters located in the same row in the convolution kernel.
[0192] The row convolution table stores the row convolution results corresponding to the word vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the row position of the convolution kernel. The row convolution result is obtained by convolving the word vector with the convolution kernel row vector corresponding to the convolution kernel identifier and the row position. The computer device can construct an index for determining the row convolution result according to the word vector of the row convolution result, the row position of the convolution kernel row vector in the convolution kernel, and the convolution kernel identifier of the convolution kernel to which the convolution kernel row vector belongs; the index can be a three-level index, in which the first-level index is used to characterize the word vector of the row convolution result, and the first-level index can be the word vector identifier; the second-level index is used to characterize the convolution kernel to which the convolution kernel row vector of the row convolution result belongs, and the second-level index can be the convolution kernel identifier; the third-level index is used to characterize the row position of the convolution kernel row vector of the row convolution result in the convolution kernel, and the third-level index can be the row position. Specifically, the structure of the three-level index is [word vector_index][convolution kernel_index][row position_index]; among them, word vector_index is the word vector identifier, word vector_index is used to represent the word vector, convolution kernel_index is the convolution kernel identifier, convolution kernel_index is used to represent the convolution kernel, row position_index is the row position, and row position_index is used to represent the position of the convolution kernel row vector in the convolution kernel (it can also be understood as representing the position of the vector in the convolution window).
[0193] The convolution kernel row vector corresponding to the convolution kernel identifier and row position refers to: a vector composed of the convolution kernel parameters at the row position in the convolution kernel corresponding to the convolution kernel identifier. For example, if the convolution kernel identifier is 2 and the row position is the second row, then the convolution kernel row vector corresponding to the convolution kernel identifier 2 and the second row is a vector composed of one or more convolution kernel parameters located in the second row of the convolution kernel of the second convolution kernel.
[0194] Each word vector belongs to a word vector set generated based on a pre-trained neural network model, that is, each representation vector is generated by a representation vector generation unit of the pre-trained neural network model.
[0195] If one of the representation vectors in the current convolution window of the target convolution kernel is the word vector token1, and the convolution kernel row vector covering token1 in the current convolution window is the convolution kernel row vector of the third row of the third convolution kernel, then the computer device can use token1 as the first-level index, the convolution kernel identifier 3 of the third convolution kernel as the second-level index, and the row position of token1 in the current convolution window as the third-level index. The obtained third-level index is exemplarily [1][3][2]. Then, based on the third-level index [1][3][2], the computer device Figure 4 The corresponding row convolution result is determined in the three-level index-row convolution table shown and used as the row convolution result of token1 and the convolution kernel row vector of the third row of the third convolution kernel.
[0196] Step S704, adding the row convolution results between each word vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel, to obtain the convolution result corresponding to the target convolution kernel in the current convolution window;
[0197] In the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input text.
[0198] like Figure 5As shown, the third convolution kernel is used as the target convolution kernel, and the word vectors in the current convolution window of the third convolution kernel are token0, token1, and token2. Among them, the row position of token0 in the current convolution window is the first row, the row position of token1 in the current convolution window is the second row, and the row position of token2 in the current convolution window is the third row. Based on the three-level index introduced in step S203, the computer device can determine the row convolution result between token0 and the first row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by the index [0][3][0]), determine the row convolution result between token1 and the second row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by the index [1][3][1]), determine the row convolution result between token2 and the third row of the convolution kernel row vector of the third convolution kernel from the row convolution table (the row convolution result can be represented by the index [2][3][2]), add the convolution results of each row, and use the result [0][3][0]+[1][3][1]+[2][3][2] obtained by the addition as the convolution result corresponding to the current convolution window of the third convolution kernel, and the convolution result is used to generate a feature representation of the input text.
[0199] In the above text processing method, the convolution operation is split into offline and online. In the offline stage, the row convolution table is first used to store the row convolution results corresponding to the word vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network, and the row position of the convolution kernel. In the online stage, in the process of the target convolution kernel sliding over the representation matrix of the input text, for each word vector located in the current convolution window of the target convolution kernel, the word vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the word vector in the current convolution window is used as the third-level index. The row convolution table constructed in the offline stage is queried to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel. The above offline stage performs the multiplication operation of the convolution to obtain the row convolution result. The online stage can obtain the row convolution result by querying the row convolution table, and the online stage only needs to perform a simple addition on the row convolution result to obtain the convolution result corresponding to the current convolution window, which improves the speed of online convolution calculation and reduces the time of online convolution calculation.
[0200] It should be noted that the specific embodiments of the above-mentioned text processing method can refer to the embodiments of the object processing method provided in this application.
[0201] In order to better understand the object processing method or text processing method provided by the present application, an application example is described in detail below.
[0202] In order to solve the problem of "high latency" in online convolution calculation of pre-trained neural networks, this application embodiment provides a method of separating calculations, transferring the larger part of the calculation amount in the convolution operation to offline processing (i.e., offline calculation), and processing the smaller part of the calculation amount in the convolution operation online, and splitting the convolution operation into two parts: offline operation and online calculation. Among them, the part of the convolution operation transferred to offline processing is the product operation between floating-point numbers, and the part processed online is the addition operation, which can reduce the time of online convolution calculation.
[0203] In this application embodiment, offline processing includes: constructing a three-level index-row convolution table, and storing the three-level index-row convolution table in a database or file, and online processing includes: loading the three-level index-row convolution table or calling the database or file storing the three-level index-row convolution table.
[0204] In the field of artificial intelligence such as image recognition, speech recognition and text processing, although the input object is random, for the machine, the set of units that make up the input object is determined, such as the set of pixels that make up the image, the set of speech frames that make up the audio, and the set of text units that make up the text (text units can be called words or phrases). Take text as an example: text units are the most fine-grained components that can represent text, such as characters and words in Chinese, and characters and words in English. In the pre-trained neural network model, the general task is to fit the parameters of the labeled training data set through machine learning methods, and use the fitted model to make predictions on the data set to be predicted.
[0205] Taking text as an example, in order to speed up the online response speed of text prediction using artificial neural networks, in the offline case, the word vector of each word is convolved with the convolution kernel row vector of the convolution kernel and the obtained row convolution results are stored, and a three-level index-row convolution table is constructed; when predicting the input text online, the three-level index-row convolution table can be searched to obtain the corresponding row convolution results, and the row convolution results can be added and processed to obtain the output result of the convolution feature extraction unit of the pre-trained neural network model.
[0206] Offline processing specifically includes: after constructing the pre-trained neural network model, convolving each word vector of the vocabulary with the convolution kernel row vector of each convolution kernel (this convolution operation belongs to pre-calculation), storing the obtained row convolution results, and storing the constructed three-level index-row convolution table in text form, or storing it in a database.
[0207] Specifically, the vocabulary includes N+1 word vectors, and each word vector is convolved with each convolution kernel row vector in the above four convolution kernels in turn (bitwise product summation). For example, the 8-bit data of token0 is x0~x7, and the 8-bit data of the convolution kernel row vector of the first row of the first convolution kernel is f00~f07. The bitwise product summation of token0 and the convolution kernel row vector of the first row of the first convolution kernel is obtained, and the result x0×f00+x1×01+x2×f02+x3×f03+x4×f04+x5×f05+x6×f06+x7×f07 is used as the row convolution result.
[0208] Convolve each word vector of the vocabulary with each convolution kernel row vector of each convolution kernel, and after obtaining the corresponding row convolution results, construct a three-level index of each row convolution result. The structure of the three-level index is [word vector_index][convolution kernel_index][row position_index]; where word vector_index is the word vector identifier, word vector_index is used to represent the word vector, convolution kernel_index is the convolution kernel identifier, convolution kernel_index is used to represent the convolution kernel, and row position_index is the row position, which is used to represent the position of the convolution kernel row vector in the convolution kernel (it can also be understood as indicating the position of the vector in the convolution window). Based on each row convolution result and the three-level index corresponding to the row convolution result, a three-level index-row convolution table is obtained and stored in the database for online calculation.
[0209] The vocabulary consisting of each word vector generated by the representation vector generation unit based on the pre-trained neural network model is as follows Figure 4 As shown, the vocabulary includes N+1 word vectors, namely token0, token1, token2, ..., tokenN, and the dimension of each word vector is 8; the convolution kernel of the convolution feature extraction unit of the pre-trained neural network model includes the first convolution kernel, the second convolution kernel, the third convolution kernel and the fourth convolution kernel, and the scale of each convolution kernel is different, and the corresponding convolution window is also different.
[0210] After the pre-trained neural network model is constructed, each word vector of the vocabulary is convolved with the convolution kernel row vector of each convolution kernel (this convolution operation belongs to pre-calculation), the row convolution results are stored, and the constructed three-level index-row convolution table is stored in text form or stored in a database.
[0211] Specifically, the vocabulary includes N+1 word vectors, and each word vector is convolved with each convolution kernel row vector in the above four convolution kernels in turn (bitwise product summation). For example, the 8-bit data of token0 is x0~x7, and the 8-bit data of the convolution kernel row vector of the first row of the first convolution kernel is f00~f07. The bitwise product summation of token0 and the convolution kernel row vector of the first row of the first convolution kernel is obtained, and the result x0×f00+x1×01+x2×f02+x3×f03+x4×f04+x5×f05+x6×f06+x7×f07 is used as the row convolution result.
[0212] Convolve each word vector of the vocabulary with each convolution kernel row vector of each convolution kernel, and after obtaining the corresponding row convolution results, construct a three-level index of each row convolution result. The structure of the three-level index is [word vector_index][convolution kernel_index][row position_index]; where word vector_index is the word vector identifier, word vector_index is used to represent the word vector, convolution kernel_index is the convolution kernel identifier, convolution kernel_index is used to represent the convolution kernel, and row position_index is the row position, which is used to represent the position of the convolution kernel row vector in the convolution kernel (it can also be understood as indicating the position of the vector in the convolution window). Based on each row convolution result and the three-level index corresponding to the row convolution result, a three-level index-row convolution table is obtained and stored in the database for online calculation.
[0213] The online processing specifically includes: segmenting the input text to obtain the words that make up the input text; for each word, by querying the word list built based on the pre-trained neural network, obtaining the representation vector of each word (hereinafter referred to as the word vector); combining and arranging the word vectors to obtain the representation matrix of the input text (the representation matrix can be two-dimensional, the length of the row direction of the representation matrix represents the number of words that make up the input text, and the length of the column direction of the representation matrix represents the dimension of the word vector); obtaining multiple convolution kernels of the convolution feature extraction unit of the pre-trained neural network, and the size of each convolution kernel can be different; performing convolution operation on each convolution kernel and the representation matrix of the input text (bitwise multiplication and then summing) to obtain the convolution result, and performing other processing such as pooling on the convolution result.
[0214] Specifically, before the input text enters the convolution feature extraction unit of the pre-trained neural network model for feature extraction, the input text will be segmented to obtain the words that make up the input text, and then the word vectors of each word are determined by looking up the word table, and the word vectors of each word are arranged to obtain the representation matrix of the input text. Assuming that the number of words that make up the input text is n, the word vectors of each word are found by looking up the word table. Among them, the word vectors of each word belong to vectors with a fixed dimension (such as dim dimension) generated by the representation vector generation unit based on the pre-trained neural network model. According to the order of each word in the input text, the word vectors of each word are arranged to obtain the n×dim representation matrix of the input text, and then the representation matrix is convolved with the convolution kernel (the convolution operation includes multiplication between data types such as floating point numbers, addition between data types such as floating point numbers, etc.), and the convolution result is passed to the subsequent layer (such as the pooling layer, etc.), until the calculation result is finally output.
[0215] like Figure 8 As shown, in the online operation part, the steps performed by the computer device include:
[0216] Step S801, performing word segmentation processing on the input text;
[0217] Step S802, obtaining the word vector of each word to obtain a representation matrix of the input text;
[0218] Step S803, according to the row position of each word vector in the convolution window, query the three-level index-row convolution table to obtain the row convolution result between each word vector and the convolution kernel row vector of the corresponding row position;
[0219] That is to say, when the convolution kernel slides on the representation matrix, the row convolution result between each word vector located in the current convolution window and the convolution kernel row vector at the corresponding position is determined by looking up the three-level index-row convolution table.
[0220] Step S804, summing the row convolution results of the same convolution window to obtain the convolution result of the same convolution window, so as to obtain the convolution result of the corresponding convolution kernel and the representation matrix of the input text;
[0221] After determining the row convolution results between each word vector in the current convolution window and the row vector of the convolution kernel at the corresponding position, the computer device sums the row convolution results to obtain the convolution result corresponding to the current convolution window. According to the sliding position, the convolution results of the same convolution kernel at each sliding position are arranged to obtain the convolution kernel and the row vector of the convolution kernel. Figure 5 The convolution result of the text representation matrix is shown to be input into the pooling layer for pooling processing.
[0222] For example, if the number of words constituting the input text is 5, such as Figure 5As shown in the figure, the word vectors of each word are token0, token1, token2, token3 and token4 respectively; each word vector is convolved with each convolution kernel row vector in the four convolution kernels in turn.
[0223] Among them, the first convolution kernel includes a convolution kernel row vector, and the first convolution kernel slides over Figure 5 After the text representation matrix is shown, there are 5 row convolution results, which are the convolution operation results between each word vector and the convolution kernel row vector of the first convolution kernel. The three-level indexes corresponding to each row convolution result are: [0][1][0], [1][1][0], [2][1][0], [3][1][0], [4][1][0].
[0224] The second convolution kernel includes two convolution kernel row vectors. The second convolution kernel slides over Figure 5 After the text representation matrix is shown, there are 8 row convolution results, and the three-level indexes corresponding to each row convolution result are: [0][2][0], [1][2][0], [2][2][0], [3][2][0], [1][2][1], [2][2][1], [3][2][1], [4][2][1]; among them, the row convolution results corresponding to the same convolution window are summed to obtain [0][2][0]+[1][2][1], [1][2][0]+[2][2][1], [2][2][0]+[3][2][1], [3][2][0]+[4][2][1].
[0225] For the third and fourth convolution kernels, respectively Figure 5 The convolution operation of the text representation matrix shown is similar to the above process and will not be repeated here.
[0226] This application embodiment is applied to a real online system, which can effectively reduce the online computing response time of convolution. Since short text computing fields such as search, recommendation, and dialogue systems involve many shallow neural network applications, the computing efficiency improved by the solution provided by this embodiment is particularly prominent. Compared with not using the solution provided by this embodiment, the online system throughput is improved by more than 10 times. The online computing time of convolution can be greatly reduced, and the online computing does not involve multiplication operations between types of data such as floating point numbers; and the solution provided by this embodiment has high portability, does not involve changes and adaptations at the computer hardware and operating system level, is very flexible, and has high popularity.
[0227] It should be understood that although Figures 2 to 8The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 to 8 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0228] In some embodiments, Fig. 9 As shown, an object processing device is provided, comprising:
[0229] A first representation matrix determination module 901 is used to determine a representation matrix of an input object based on a pre-trained neural network model;
[0230] A first convolution kernel sliding module 902 is used to determine each representation vector located in a current convolution window from the representation matrix in a process in which the target convolution kernel slides over the representation matrix based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model;
[0231] The first result query module 903 is used to query the row convolution table generated based on the pre-trained neural network model for each representation vector located in the current convolution window, using the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel;
[0232] The first result processing module 904 is used to add the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input object.
[0233] In some embodiments, the first representation matrix determination module 901 is used to divide the input object into multiple units; obtain a representation vector set generated based on the pre-trained neural network model; each representation vector in the representation vector set is used to represent each unit that constitutes the object; the representation vectors in the representation vector set are obtained by inputting the units that constitute the object into the pre-trained neural network model; for each unit that constitutes the input object, query the representation vector used to represent the unit from the representation vector set; and determine the representation matrix of the input object based on the representation vector corresponding to each unit.
[0234] In some embodiments, the first representation matrix determination module 901 is further used to determine the distribution of each unit constituting the input object in the input object; and to combine the representation vectors corresponding to each unit according to the distribution to obtain the representation matrix of the input object.
[0235] In some embodiments, the apparatus further comprises: a representation vector representation determination module, configured to determine a representation vector identifier of each representation vector in the representation matrix of the input object;
[0236] The first result query module 903 is also used to query the row convolution table generated based on the pre-trained neural network model for each representation vector located in the current convolution window, using the representation vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel.
[0237] In some embodiments, the input object is input text, and each representation vector in the representation vector set is used to represent each word constituting the text, and the representation vector is a word vector obtained by inputting the word into the pre-trained neural network model for processing text; or, the input object is audio, and each representation vector in the representation vector set is used to represent each segment constituting the audio, and the representation vector is an audio vector obtained by inputting the segment into the pre-trained neural network model for processing audio.
[0238] In some embodiments, the device also includes: a row convolution table construction module, which is used to input the various units that make up the sample object into the pre-trained neural network model, obtain the representation vector corresponding to each unit, and generate a unit identifier for each unit or a vector identifier corresponding to each unit; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the representation vector; convolve each representation vector with each convolution kernel row vector to obtain a row convolution result; use the unit identifier or vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index; store each of the third-level index in correspondence with the corresponding row convolution result to generate a row convolution table.
[0239] In some embodiments, the row convolution table construction module is also used to traverse the representation vector and the convolution kernel in the first convolution layer of the pre-trained neural network model; determine the representation vector and the convolution kernel row vector to be convolved; multiply the elements corresponding to the same position in the representation vector and the convolution kernel row vector to obtain the product results of each position; sum the product results of each position to obtain the row convolution result between the representation vector to be convolved and the convolution kernel row vector.
[0240] In some embodiments, the object is text, and a row convolution table construction module is used to input each word constituting the sample text into the pre-trained neural network model, obtain the word vector corresponding to each word, and generate a word identifier or a corresponding word vector identifier for each word; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the word vector; convolve each word vector with each convolution kernel row vector to obtain a row convolution result; use the word identifier or word vector identifier corresponding to the word vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index; store each of the third-level index in correspondence with the corresponding row convolution result to generate a row convolution table.
[0241] In some embodiments, the device also includes: a convolution result determination module, used to determine the current sliding position of the current convolution window in the representation matrix, and use the convolution result corresponding to the current convolution window as the convolution result corresponding to the current sliding position; based on the convolution result corresponding to each sliding position in the process of the target convolution kernel sliding across the representation matrix, generate the convolution result of the representation matrix and the target convolution kernel.
[0242] In some embodiments, the device also includes: a pooling processing module, which is used to perform pooling processing on the convolution result of the representation matrix and the target convolution kernel through a pooling layer in the pre-trained neural network to obtain a pooling result.
[0243] In some embodiments, when the object is an image, the device further includes: an image convolution processing module, which is used to obtain a pixel matrix of an input image; based on the scale of a target convolution kernel to be convolved with the pixel matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding over the pixel matrix, determining each pixel value located in a current convolution window from the pixel matrix; for each pixel value located in the current convolution window, using the pixel value as a first-level index, the convolution kernel identifier of the target convolution kernel as a second-level index, and the position of the pixel value in the current convolution window as a third-level index, querying a product table generated based on the pre-trained neural network model to obtain the pixel value and the target convolution kernel. The product result between the convolution kernel parameters at the position in the convolution kernel; adding the product results between each pixel value in the current convolution window and the convolution kernel parameters at the corresponding position in the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein the product table stores the product result corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model and the position of the convolution kernel parameters, and the product result is obtained by multiplying the pixel value with the convolution kernel parameters at the position in the convolution kernel identified by the convolution kernel identifier; in the process of the target convolution kernel sliding across the pixel matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input image.
[0244] In the above-mentioned object processing device, the convolution operation is split into offline and online. In the offline stage, the row convolution table is first used to store the row convolution result corresponding to the representation vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network, and the row position of the convolution kernel. In the online stage, in the process of the target convolution kernel sliding over the representation matrix of the input object, for each representation vector located in the current convolution window of the target convolution kernel, the row convolution table constructed in the offline stage is queried with the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain the row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel. The above-mentioned offline stage performs the product operation of convolution to obtain the row convolution result, and the online stage can obtain the row convolution result by querying the row convolution table, and the online stage only needs to perform a simple addition on the row convolution result to obtain the convolution result corresponding to the current convolution window, thereby improving the speed of online convolution calculation and reducing the time of online convolution calculation.
[0245] For the specific definition of the object processing device, please refer to the definition of the object processing method above, which will not be repeated here. Each module in the above-mentioned object processing device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0246] In some embodiments, Fig.10 As shown, a text processing device is provided, comprising:
[0247] The second representation matrix determination module 1001 is used to determine the representation matrix of the input text based on the pre-trained neural network model; the representation matrix of the input text is obtained according to the word vector of each word in the input text;
[0248] A second convolution kernel sliding module 1002 is used to determine each word vector located in a current convolution window from the representation matrix in a process in which the target convolution kernel slides over the representation matrix based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model;
[0249] The second result query module 1003 is used to query the row convolution table generated based on the pre-trained neural network model for each word vector located in the current convolution window, using the word vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the word vector in the current convolution window as the third-level index, to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel;
[0250] The second result processing module 1004 is used to add the row convolution results between each word vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input text.
[0251] In the above-mentioned text processing device, the convolution operation is split into offline and online. In the offline stage, the row convolution table is first used to store the row convolution results corresponding to the word vector, the convolution kernel identifier of the convolution kernel in the pre-trained neural network, and the row position of the convolution kernel. In the online stage, in the process of the target convolution kernel sliding over the representation matrix of the input text, for each word vector located in the current convolution window of the target convolution kernel, the word vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the word vector in the current convolution window is used as the third-level index. The row convolution table constructed in the offline stage is queried to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel. The above-mentioned offline stage performs a multiplication operation of the convolution to obtain the row convolution result. The online stage can obtain the row convolution result by querying the row convolution table, and the online stage only needs to do a simple addition of the row convolution result to obtain the convolution result corresponding to the current convolution window, which improves the speed of online convolution calculation and reduces the time of online convolution calculation.
[0252] For the specific definition of the text processing device, please refer to the definition of the text processing method above, which will not be repeated here. Each module in the above text processing device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0253] In one embodiment, a computer device is provided, which may be Figure 1 The terminal 102 or server 104 shown in the figure may have an internal structure as shown in FIG. Fig.11As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data of the object processing method or the text processing method. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer device also includes an input and output interface, which is a connection circuit for exchanging information between the processor and the external device, and they are connected to the processor through a bus, referred to as an I / O interface. When the computer program is executed by the processor, one of the above-mentioned object processing methods or the above-mentioned text processing methods is implemented.
[0254] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0255] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0256] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0257] In some embodiments, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments are implemented.
[0258] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0259] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the above-mentioned computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0260] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0261] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. An object processing method, characterized in that: The method comprises: Determine the representation matrix of the input object based on the pre-trained neural network model; Based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding across the representation matrix, determining each representation vector located in the current convolution window from the representation matrix; For each representation vector located in the current convolution window, the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index, and a row convolution table generated based on the pre-trained neural network model is queried to obtain a row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel; Add the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input object.
2. The method according to claim 1, characterized in that The step of determining a representation matrix of an input object based on a pre-trained neural network model includes: Splitting the input object into a plurality of units; Acquire a set of representation vectors generated based on the pre-trained neural network model; each representation vector in the set of representation vectors is used to represent each unit constituting an object; the representation vectors in the set of representation vectors are obtained by inputting the units constituting the object into the pre-trained neural network model; For each unit constituting the input object, querying a representation vector for characterizing the unit from the representation vector set; A representation matrix of the input object is determined according to the representation vector corresponding to each unit.
3. The method according to claim 2, characterized in that Determining the representation matrix of the input object according to the representation vector corresponding to each unit includes: Determining the distribution of the various units constituting the input object in the input object; The representation vectors corresponding to each unit are combined according to the distribution to obtain a representation matrix of the input object.
4. The method according to claim 2, characterized in that: The method further comprises: Determining a representation vector identifier of each representation vector in a representation matrix of the input object; For each representation vector located in the current convolution window, the representation vector is used as a first-level index, the convolution kernel identifier of the target convolution kernel is used as a second-level index, and the row position of the representation vector in the current convolution window is used as a third-level index, and a row convolution table generated based on the pre-trained neural network model is queried to obtain a row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel, including: For each representation vector located in the current convolution window, the representation vector identifier corresponding to the representation vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the representation vector in the current convolution window is used as the third-level index. The row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the representation vector and the row vector at the corresponding row position of the target convolution kernel.
5. The method according to claim 2, characterized in that: The input object is an input text, each representation vector in the representation vector set is used to represent each word constituting the text, and the representation vector is a word vector obtained by inputting the word into the pre-trained neural network model for processing the text; or, The input object is audio, and each representation vector in the representation vector set is used to characterize each segment constituting the audio. The representation vector is an audio vector obtained by inputting the segment into the pre-trained neural network model for processing audio.
6. The method according to claim 2, characterized in that The steps of constructing the row convolution table include: Input each unit constituting the sample object into the pre-trained neural network model, obtain a representation vector corresponding to each unit, and generate a unit identifier of each unit or a vector identifier corresponding to each unit; For each convolution kernel in the first convolution layer of the pre-trained neural network model, determine a convolution kernel identifier of each convolution kernel and a convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the representation vector; Convolve each representation vector with each convolution kernel row vector to obtain the row convolution result; A third-level index is generated by taking the unit identifier or vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index; Each of the three-level indexes is stored corresponding to the corresponding row convolution result to generate a row convolution table.
7. The method according to claim 6, characterized in that The convolution operation is performed on each representation vector and each convolution kernel row vector to obtain a row convolution result, including: Traversing the representation vector and the convolution kernel in the first convolution layer of the pre-trained neural network model; Determine a representation vector to be convolutionally operated and a row vector of the convolution kernel; Multiplying the elements at the same position corresponding to the representation vector and the convolution kernel row vector to obtain a product result at each position; The product results of each position are summed to obtain a row convolution result between the representation vector to be convolved and the convolution kernel row vector.
8. The method according to claim 1, characterized in that The object is text, and the steps of constructing the row convolution table include: Input each word constituting the sample text into the pre-trained neural network model to obtain a word vector corresponding to each word, and generate a word identifier or a corresponding word vector identifier for each word; For each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the word vector; Convolve each word vector with each convolution kernel row vector to obtain the row convolution result; A third-level index is generated by taking the word identifier or word vector identifier corresponding to the word vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index; Each of the three-level indexes is stored corresponding to the corresponding row convolution result to generate a row convolution table.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Determine a current sliding position of the current convolution window in the representation matrix, and use the convolution result corresponding to the current convolution window as the convolution result corresponding to the current sliding position; According to the convolution result corresponding to each sliding position in the process of the target convolution kernel sliding across the representation matrix, the convolution result of the representation matrix and the target convolution kernel is generated.
10. The method according to claim 9, characterized in that The method further comprises: The convolution result of the representation matrix and the target convolution kernel is pooled through the pooling layer in the pre-trained neural network to obtain a pooling result.
11. The method according to claim 1, characterized in that When the object is an image, the method further includes: Get the pixel matrix of the input image; Based on the scale of the target convolution kernel to be convolved with the pixel matrix in the pre-trained neural network model, determine each pixel value located in the current convolution window from the pixel matrix in the process of the target convolution kernel sliding across the pixel matrix; For each pixel value located in the current convolution window, the pixel value is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the position of the pixel value in the current convolution window is used as the third-level index, and a product table generated based on the pre-trained neural network model is queried to obtain a product result between the pixel value and the convolution kernel parameter at the position in the target convolution kernel; Add the product results of each pixel value in the current convolution window and the convolution kernel parameters at the corresponding position in the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; Among them, the product table stores the product results corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the position of the convolution kernel parameters, and the product result is obtained by multiplying the pixel value with the convolution kernel parameter at the position in the convolution kernel identified by the convolution kernel identifier; in the process of the target convolution kernel sliding across the pixel matrix, the convolution results corresponding to each convolution window are used to generate a feature representation of the input image.
12. A text processing method, characterized in that: The method comprises: Determine a representation matrix of an input text based on a pre-trained neural network model; the representation matrix of the input text is obtained according to a word vector of each word in the input text; Based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, determining each word vector located in the current convolution window from the representation matrix in the process of the target convolution kernel sliding across the representation matrix; For each word vector located in the current convolution window, the word vector is used as the first-level index, the convolution kernel identifier of the target convolution kernel is used as the second-level index, and the row position of the word vector in the current convolution window is used as the third-level index. The row convolution table generated based on the pre-trained neural network model is queried to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel; Add the row convolution results between each word vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input text.
13. An object processing device, characterized in that: The device comprises: A first representation matrix determination module, used to determine a representation matrix of an input object based on a pre-trained neural network model; A first convolution kernel sliding module, configured to determine, based on a scale of a target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model, each representation vector located within a current convolution window from the representation matrix in a process in which the target convolution kernel slides across the representation matrix; A first result query module is used to query the row convolution table generated based on the pre-trained neural network model for each representation vector located in the current convolution window, using the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain a row convolution result between the representation vector and the row vector of the corresponding row position of the target convolution kernel; The first result processing module is used to add the row convolution results between each representation vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input object.
14. The device according to claim 13, characterized in that The first representation matrix determination module is used to divide the input object into multiple units; obtain a representation vector set generated based on the pre-trained neural network model; each representation vector in the representation vector set is used to represent each unit that constitutes the object; the representation vectors in the representation vector set are obtained by inputting the units that constitute the object into the pre-trained neural network model; for each unit that constitutes the input object, query the representation vector used to represent the unit from the representation vector set; and determine the representation matrix of the input object based on the representation vector corresponding to each unit.
15. The device according to claim 14, characterized in that The first representation matrix determination module is further used to determine the distribution of each unit constituting the input object in the input object; and to combine the representation vectors corresponding to each unit according to the distribution to obtain the representation matrix of the input object.
16. The device according to claim 14, characterized in that The device further comprises: a representation vector representation determination module, configured to determine a representation vector identifier of each representation vector in the representation matrix of the input object; The first result query module is also used to query the row convolution table generated based on the pre-trained neural network model for each representation vector located in the current convolution window, using the representation vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the representation vector in the current convolution window as the third-level index, to obtain the row convolution result between the representation vector and the row vector at the corresponding row position of the target convolution kernel.
17. The device according to claim 14, characterized in that The input object is input text, and each representation vector in the representation vector set is used to represent each word constituting the text, and the representation vector is a word vector obtained by inputting the word into the pre-trained neural network model for processing text; or, the input object is audio, and each representation vector in the representation vector set is used to represent each segment constituting the audio, and the representation vector is an audio vector obtained by inputting the segment into the pre-trained neural network model for processing audio.
18. The device according to claim 14, characterized in that The device also includes: a row convolution table construction module, which is used to input the various units that constitute the sample object into the pre-trained neural network model, obtain the representation vector corresponding to each unit, and generate a unit identifier of each unit or a vector identifier corresponding to each unit; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the representation vector; perform convolution operations on each representation vector and each convolution kernel row vector respectively to obtain a row convolution result; use the unit identifier or vector identifier corresponding to the representation vector as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel as the third-level index to generate a third-level index; store each of the third-level indexes in correspondence with the corresponding row convolution result to generate a row convolution table.
19. The device according to claim 18, characterized in that The row convolution table construction module is also used to traverse the representation vector and the convolution kernel in the first convolution layer of the pre-trained neural network model; determine the representation vector to be convolved and the convolution kernel row vector; multiply the elements corresponding to the same position in the representation vector and the convolution kernel row vector to obtain the product results of each position; sum the product results of each position to obtain the row convolution result between the representation vector to be convolved and the convolution kernel row vector.
20. The device according to claim 13, characterized in that The object is text, and the row convolution table construction module is used to input each word constituting the sample text into the pre-trained neural network model, obtain the word vector corresponding to each word, and generate a word identifier or a corresponding word vector identifier for each word; for each convolution kernel in the first convolution layer of the pre-trained neural network model, determine the convolution kernel identifier of each convolution kernel and the convolution kernel row vector of each convolution kernel; the dimension of the convolution kernel row vector is the same as the dimension of the word vector; convolution operation is performed on each word vector with each convolution kernel row vector to obtain a row convolution result; the word identifier or word vector identifier corresponding to the word vector is used as the first-level index, the convolution kernel identifier corresponding to the convolution kernel where the convolution kernel row vector is located is used as the second-level index, and the row position of the convolution kernel row vector in the convolution kernel is used as the third-level index to generate a third-level index; each of the third-level index is stored in correspondence with the corresponding row convolution result to generate a row convolution table.
21. The device according to any one of claims 13 to 20, characterized in that The device also includes: a convolution result determination module, which is used to determine the current sliding position of the current convolution window in the representation matrix, and use the convolution result corresponding to the current convolution window as the convolution result corresponding to the current sliding position; according to the convolution result corresponding to each sliding position in the process of the target convolution kernel sliding across the representation matrix, the convolution result of the representation matrix and the target convolution kernel is generated.
22. The device according to claim 21, characterized in that The device also includes: a pooling processing module, which is used to perform pooling processing on the convolution result of the representation matrix and the target convolution kernel through the pooling layer in the pre-trained neural network to obtain a pooling result.
23. The device according to claim 13, characterized in that When the object is an image, the device also includes: an image convolution processing module, which is used to obtain a pixel matrix of an input image; based on the scale of a target convolution kernel to be convolved with the pixel matrix in the pre-trained neural network model, in the process of the target convolution kernel sliding over the pixel matrix, determining each pixel value located in a current convolution window from the pixel matrix; for each pixel value located in the current convolution window, using the pixel value as a first-level index, the convolution kernel identifier of the target convolution kernel as a second-level index, and the position of the pixel value in the current convolution window as a third-level index, querying a product table generated based on the pre-trained neural network model to obtain the pixel value and the pixel value in the target convolution kernel. The product result between the convolution kernel parameters at the position is added; the product result between each pixel value in the current convolution window and the convolution kernel parameters at the corresponding position in the target convolution kernel is added to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein the product table stores the product result corresponding to the pixel value, the convolution kernel identifier of the convolution kernel in the pre-trained neural network model, and the position of the convolution kernel parameters, and the product result is obtained by multiplying the pixel value with the convolution kernel parameters at the position in the convolution kernel identified by the convolution kernel identifier; in the process of the target convolution kernel sliding across the pixel matrix, the convolution results corresponding to each convolution window are used to generate a feature representation of the input image.
24. A text processing device, characterized in that: The device comprises: A second representation matrix determination module is used to determine a representation matrix of an input text based on a pre-trained neural network model; the representation matrix of the input text is obtained according to a word vector of each word in the input text; A second convolution kernel sliding module is used to determine each word vector located in a current convolution window from the representation matrix in a process in which the target convolution kernel slides across the representation matrix based on the scale of the target convolution kernel to be convolved with the representation matrix in the pre-trained neural network model; A second result query module is used to query the row convolution table generated based on the pre-trained neural network model for each word vector located in the current convolution window, using the word vector as the first-level index, the convolution kernel identifier of the target convolution kernel as the second-level index, and the row position of the word vector in the current convolution window as the third-level index, to obtain the row convolution result between the word vector and the row vector of the corresponding row position of the target convolution kernel; The second result processing module is used to add the row convolution results between each word vector in the current convolution window and the row vector at the corresponding row position of the target convolution kernel to obtain the convolution result corresponding to the target convolution kernel in the current convolution window; wherein, in the process of the target convolution kernel sliding across the representation matrix, the convolution results corresponding to each convolution window are used to generate the feature representation of the input text.
25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.
26. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
27. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Data processing method, neural network training method and related device and apparatus
CN113326930A
Method and apparatus for complexity reduction on two-dimensional convolutions for image processing
US6151025A