Data processing method, model training method and related equipment
By updating the image feature information of the neural network layer weight parameters of the machine learning model, the problem of high computational resource consumption in multimodal data processing is solved, and efficient and accurate multimodal data processing is achieved.
Patent Information
- Application Number
- CN202410547007.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing machine learning models mainly process single-modal data and have difficulty efficiently processing multimodal data, resulting in excessive consumption of computing resources.
By extracting image feature information, the weight parameters of the neural network layers in the machine learning model are updated, and the model is optimized using image information without changing the input length of text and speech, thereby reducing the consumption of additional computing resources.
It effectively processes multimodal data, improves the accuracy of prediction information, reduces the consumption of computing resources, and maintains the fluency and efficiency of the model.
Smart Images

Figure CN120876882A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a data processing method, a model training method, and related equipment. Background Technology
[0002] With the development of artificial intelligence technology, machine learning models are being applied to more scenarios. In some scenarios, it is necessary to process multimodal data. For example, multimodal data can include both images and text, or both images and speech. However, most current machine learning models are used to process single-modal data, such as pure images, pure text, or pure speech. Therefore, a solution that can process multimodal data is urgently needed. Summary of the Invention
[0003] This application provides a data processing method, a model training method, and related equipment. It provides a solution capable of processing multimodal data and uses image feature information to update the weight parameters of the neural network layer in the machine learning model without changing the length of the text and / or speech input into the machine learning model, thereby minimizing the additional computer resource consumption caused by processing multimodal data.
[0004] The embodiments of this application provide the following technical solutions:
[0005] In a first aspect, embodiments of this application provide a data processing method that can use artificial intelligence technology to process multimodal data. In this method, after the execution device acquires first data (the "first data" in this application can also be understood as "multimodal data"), the first data includes images and also includes text and / or voice. Features can be extracted from the images in the first data to obtain first feature information of the images. Based on the first feature information, the weight parameters of at least one first neural network layer in the machine learning model are updated to obtain an updated machine learning model, and then the prediction information corresponding to the first data generated by the updated machine learning model is obtained.
[0006] The updated machine learning model includes updated weight parameters for each of the first neural network layers in at least one first neural network layer. The input of the updated machine learning model includes text and / or speech, or it can also be referred to as the input of the machine learning model including text and / or speech, that is, the machine learning model is a model for processing text and / or speech.
[0007] The prediction information corresponding to the first data is generated by the updated machine learning model. For example, the prediction information corresponding to the first data can be text information or voice information.
[0008] In this implementation, after acquiring the first multimodal data, feature extraction can be performed on the image to obtain the first feature information of the image. The weight parameters of at least one first neural network layer in the machine learning model are updated using the first feature information to obtain an updated machine learning model. The input of the machine learning model includes text and / or speech. The updated machine learning model generates prediction information corresponding to the first data. Thus, in the process of generating prediction information, not only image information but also text and / or speech information are used, which provides a solution capable of processing multimodal data. In addition, in this application, the first feature information of the image is used to update the weight parameters of at least one first neural network layer in the aforementioned machine learning model. This method does not change the length of the text and / or speech input into the machine learning model, so as to minimize the additional computer resource consumption caused by processing multimodal data.
[0009] In one possible implementation, the first feature information may include C feature maps, where C is an integer greater than or equal to 1, and each feature map contains n location points, where n is an integer greater than or equal to 1. The method may further include: the execution device acquiring the location encoding information of each of the n location points. Then, based on the first feature information, the execution device updates the weight parameters of at least one first neural network layer in the machine learning model. Specifically, this may include: the execution device updating the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point.
[0010] For example, the "location encoding information of each location point" in this application means that the location encoding information of the location point is related to the location of the location point. Optionally, the location encoding information of each location point can be learned during the training phase, and the location encoding information of different location points is different; or, the location encoding information of each location point can also be obtained using a preset location encoding algorithm.
[0011] In this implementation, when updating the weight parameters of at least one first neural network layer, not only is the first feature information in the form of feature map utilized, but also the position encoding information of each location point is obtained. This allows for the use of more information obtained from the image to jointly update the weight parameters of at least one first neural network layer, which is beneficial for making fuller use of the information obtained from the image. This, in turn, helps to obtain prediction information that is more compatible with the first data, thereby improving the accuracy of the obtained prediction information.
[0012] In one possible implementation, at least one first neural network layer is a neural network layer in a first feedforward neural network (FFN) module included in the machine learning model, that is, the execution device updates the weight parameters of the first neural network layer in at least one first FFN; wherein, at least one first FFN may include all FFNs in the first machine learning model, or at least one first FFN may also include some FFNs in the first machine learning model.
[0013] In this implementation, since the FFN module generally only includes fully connected layers, the construction of the FFN module is relatively simple. Choosing to update the weight parameters of the first neural network layer in the FFN helps to reduce the complexity of the process of "updating the weight parameters of at least one first neural network layer" and provides a simpler solution.
[0014] In one possible implementation, the first neural network layer is a fully connected layer in the first FFN module of the machine learning model; a "fully connected layer" can also be called a "linear layer." In this implementation, since processing multimodal data requires fusing information from different modalities compared to processing only single-modal data, the fusion process incurs additional computer resource consumption (e.g., storage and computational resources). Because the number of parameters in the weight parameters of the fully connected layer is relatively large, the updated weight parameters of the fully connected layer are not significantly different from the original weight parameters. Therefore, the impact of using the first feature information of the image to update the weight parameters of the fully connected layer is relatively small, thereby minimizing the additional computer resource consumption during multimodal data fusion.
[0015] In one possible implementation, the execution device updates the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point. This can include: the execution device converting the first feature information into second feature information through a first neural network module, and then updating the weight parameters of at least one first neural network layer in the machine learning model based on the second feature information and the location encoding information of each location point. The second feature information can be in the form of a token, meaning its form corresponds to that of text or speech feature information.
[0016] Since the first feature information is in the form of a feature map, and in this implementation, the information obtained from the image is integrated into the processing of the first machine learning model used to process text and / or speech, the first feature information is converted into second feature information in the form of a token. The second feature information in the form of a token is consistent with the form of the feature information of text and / or speech. Then, the weight parameters of at least one first neural network layer are updated using the second feature information in the form of a token and the position encoding information of each location point. This helps to reduce the difficulty of the process of "updating the weight parameters of at least one first neural network layer", improve the degree of fusion between multimodal data, and help to obtain prediction information that is more suitable for the first data, that is, to improve the accuracy of the obtained prediction information.
[0017] In one possible implementation, the second feature information includes first sub-feature information for each location point, and the location encoding information for each location point includes first location encoding information and second location encoding information. The execution device updates the weight parameters of at least one first neural network layer in the machine learning model based on the second feature information and the location encoding information for each location point. Specifically, this may include: the execution device fusing the first sub-feature information and the first location encoding information for each location point to obtain second sub-feature information for each location point; fusing the first sub-feature information and the second location encoding information for each location point to obtain third sub-feature information for each location point; then, the execution device updates the weight parameters of the first fully connected layer using the second sub-feature information for each location point to obtain updated weight parameters of the first fully connected layer; and finally, the execution device updates the weight parameters of the second fully connected layer using the third sub-feature information for each location point to obtain updated weight parameters of the second fully connected layer.
[0018] The machine learning model includes at least one first FFN module, the first FFN module includes a first fully connected layer and a second fully connected layer, and the updated weight parameters of at least one first neural network layer include the updated weight parameters of each first FFN.
[0019] In this implementation, since the FFN module generally has a relatively simple structure, the positional encoding information of each position point includes first positional encoding information and second positional encoding information. The first sub-feature information of each position point is fused with the first positional encoding information and the second positional encoding information to obtain the second sub-feature information and the third sub-feature information of each position point. Then, the weight parameters of the first fully connected layer and the second fully connected layer are updated using the second sub-feature information and the third sub-feature information of each position point, respectively. Since updating the weight parameters of the first neural network layer may cause changes in the size of the weight parameters of the first neural network layer, which may lead to changes in the size of the information generated by the first neural network layer, choosing to update the weight parameters of the first fully connected layer and the second fully connected layer in the FFN is beneficial to achieve mutual cancellation between the influence of the first fully connected layer on the size of the processed information and the influence of the second fully connected layer on the size of the processed information. This is beneficial to ensure that the size of the information generated by the FFN module remains unchanged, thereby ensuring the smoothness of the updated first machine learning model in the data processing process.
[0020] In one possible implementation, the execution device updates the weight parameters of the first fully connected layer using the second sub-feature information of each location point, including: concatenating the weight parameters of the first fully connected layer with the second sub-feature information of each location point; the execution device updates the weight parameters of the second fully connected layer using the third sub-feature information of each location point, including: concatenating the weight parameters of the second fully connected layer with the third sub-feature information of each location point.
[0021] In this implementation, the weight parameters of the first fully connected layer are updated by concatenating them with the second sub-feature information of each location point. Similarly, the weight parameters of the second fully connected layer are updated by concatenating them with the third sub-feature information of each location point. This preserves the weight parameters of both the first and second fully connected layers before the aforementioned update. Furthermore, information obtained from the image is incorporated into the updated weight parameters of both the first and second fully connected layers. Since the weight parameters of the first and second fully connected layers were trained before the update, the accuracy of the prediction information obtained when processing text and / or speech using the updated weight parameters is high. When updating the weight parameters of both the first and second fully connected layers using information obtained from the image, the original weight parameters are preserved, which helps maintain the high accuracy of the prediction information generated by the updated first machine learning model.
[0022] In one possible implementation, the second feature information includes first sub-feature information for each location point, and the location encoding information for each location point includes first location encoding information and second location encoding information. The weight parameters of the first fully connected layer can be represented as a matrix of a first size, and the weight parameters of the second fully connected layer can be represented as a matrix of a second size. The execution device updates the weight parameters of at least one first neural network layer in the machine learning model based on the second feature information and the location encoding information for each location point. Specifically, this may include: the execution device fusing the first sub-feature information and the first location encoding information for each location point to obtain second sub-feature information for each location point; and fusing the first sub-feature information and the second location encoding information for each location point to obtain third sub-feature information for each location point.
[0023] After obtaining the second sub-feature information of each of the n location points, the execution device can place the second sub-feature information of each of the n location points into a first matrix of a first size. The empty values in the first matrix can be filled with 0. Then, the weight parameters of the first fully connected layer are added to the first matrix to obtain the updated weight parameters of the first fully connected layer.
[0024] After the execution device obtains the third sub-feature information of each of the n location points, it can place the third sub-feature information of each of the n location points into a second matrix of a second size. The empty values in the second matrix can be filled with 0. Then, the weight parameters of the second fully connected layer are added to the second matrix to obtain the updated weight parameters of the second fully connected layer.
[0025] In one possible implementation, the execution device fuses the first sub-feature information and the first position encoding information of each location point, including: adding the first sub-feature information and the first position encoding information of each location point; the execution device also fuses the first sub-feature information and the second position encoding information of each location point, including: adding the second sub-feature information and the second position encoding information of each location point. For example, the aforementioned "addition" can also be replaced by "multiplication," "subtraction," or other fusion operations.
[0026] This implementation provides a specific scheme for fusing the first sub-feature information and the first position encoding information (or second position encoding information) of each location point, which improves the feasibility of the scheme. The implementation scheme is relatively simple, thereby reducing the difficulty of fusing data from different modalities and helping to reduce the additional computer resource consumption caused by fusing data from different modalities.
[0027] In one possible implementation, the parameters that need to be adjusted during the training phase include: the weight parameters of the first neural network module and / or the positional encoding information of each location point. In this implementation, adjusting the weight parameters of the first neural network module and / or the positional encoding information of each location point during the training phase reduces the number of parameters that need to be adjusted during training. This not only helps reduce the storage resources consumed during training but also improves the efficiency of the training phase.
[0028] Secondly, embodiments of this application provide a model training method that can use artificial intelligence technology to process multimodal data. In this method, a training device acquires first data, which includes an image and further includes text and / or speech; performs feature extraction on the image to obtain first feature information of the image; updates the weight parameters of at least one first neural network layer in a machine learning model based on the first feature information to obtain an updated machine learning model, which includes updated weight parameters of at least one first neural network layer, and the input of the updated machine learning model includes text and / or speech; acquires prediction information corresponding to the first data, which is generated by the updated machine learning model; and trains the model based on expected information and prediction information corresponding to the first data, with the training objective including improving the similarity between the prediction information and the expected information.
[0029] For example, the training device trains based on expected information and predicted information corresponding to the first data, including: the training device generates a function value of a loss function based on the expected information and predicted information corresponding to the first data, and trains based on the function value of the loss function; wherein the loss function can represent the similarity between the expected information and predicted information corresponding to the first data.
[0030] In one possible implementation, the first feature information is a feature map containing at least one location point. The method further includes: a training device acquiring location encoding information for each location point among the at least one location point. The training device updates the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information, including: updating the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point.
[0031] In one possible implementation, the training device updates the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point, including: the training device converts the first feature information into second feature information through the first neural network module, the second feature information being feature information in the form of a token; and updates the weight parameters of at least one first neural network layer in the machine learning model based on the second feature information and the location encoding information of each location point.
[0032] In one possible implementation, the training device trains based on expected and predicted information corresponding to the first data, including: the training device updates the first parameters based on the expected and predicted information corresponding to the first data, the first parameters including the weight parameters of the first neural network module and / or the position encoding information of each location point.
[0033] The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the second aspect of this application can all be found in the first aspect, and will not be repeated here.
[0034] Thirdly, this application provides a data processing apparatus that can use artificial intelligence technology to process multimodal data. The apparatus includes: an acquisition module for acquiring first data, the first data including an image, and the first data further including text and / or speech; a feature extraction module for extracting features from the image to obtain first feature information of the image; an update module for updating the weight parameters of at least one first neural network layer in a machine learning model based on the first feature information to obtain an updated machine learning model, the updated machine learning model including the updated weight parameters of at least one first neural network layer, and the input of the updated machine learning model including text and / or speech; and the acquisition module is further used to acquire prediction information corresponding to the first data, the prediction information being generated by the updated machine learning model.
[0035] In one possible implementation, the first feature information is a feature map, which contains at least one location point. The acquisition module is further used to acquire the location encoding information of each location point in the at least one location point. The update module is specifically used to update the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point.
[0036] In one possible implementation, at least one first neural network layer is a neural network layer in a first feedforward neural network (FFN) module included in the machine learning model.
[0037] In one possible implementation, the first neural network layer is a fully connected layer in the first FFN module included in the machine learning model.
[0038] In one possible implementation, the update module is specifically used to: convert the first feature information into second feature information through the first neural network module, wherein the second feature information is feature information in the form of a token; and update the weight parameters of at least one first neural network layer in the machine learning model according to the second feature information and the location encoding information of each location point.
[0039] In one possible implementation, the second feature information includes first sub-feature information for each location point, and the location encoding information for each location point includes first location encoding information and second location encoding information. The update module is specifically used to: fuse the first sub-feature information and the first location encoding information for each location point to obtain the second sub-feature information for each location point; fuse the first sub-feature information and the second location encoding information for each location point to obtain the third sub-feature information for each location point; update the weight parameters of the first fully connected layer using the second sub-feature information for each location point to obtain the updated weight parameters of the first fully connected layer; and update the weight parameters of the second fully connected layer using the third sub-feature information for each location point to obtain the updated weight parameters of the second fully connected layer. The machine learning model includes at least one first FFN module, the first FFN module includes a first fully connected layer and a second fully connected layer, and the updated weight parameters of at least one first neural network layer include the updated weight parameters of each first FFN.
[0040] In one possible implementation, the update module is specifically used to: concatenate the weight parameters of the first fully connected layer with the second sub-feature information of each location point; and concatenate the weight parameters of the second fully connected layer with the third sub-feature information of each location point.
[0041] In one possible implementation, the update module is specifically used to: add the first sub-feature information of each location point to the first location encoding information of each location point; and add the second sub-feature information of each location point to the second location encoding information of each location point.
[0042] In one possible implementation, the parameters that need to be adjusted during the training phase include: the weight parameters of the first neural network module and / or the location encoding information of each location point.
[0043] The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the third aspect of this application can all be found in the first aspect, and will not be repeated here.
[0044] Fourthly, this application provides a training apparatus for a model that can use artificial intelligence technology to process multimodal data. The apparatus includes: an acquisition module for acquiring first data, the first data including an image, and the first data further including text and / or speech; a feature extraction module for extracting features from the image to obtain first feature information of the image; an update module for updating the weight parameters of at least one first neural network layer in a machine learning model according to the first feature information to obtain an updated machine learning model, the updated machine learning model including the updated weight parameters of at least one first neural network layer, and the input of the updated machine learning model including text and / or speech; the acquisition module is further used to acquire prediction information corresponding to the first data, the prediction information being generated by the updated machine learning model; and a training module for training according to expected information and prediction information corresponding to the first data, the training objective including improving the similarity between the prediction information and the expected information.
[0045] In one possible implementation, the first feature information is a feature map, which contains at least one location point. The acquisition module is further used to acquire the location encoding information of each location point in the at least one location point. The update module is specifically used to update the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point.
[0046] In one possible implementation, the update module is specifically used to: convert the first feature information into second feature information through the first neural network module, wherein the second feature information is feature information in the form of a token; and update the weight parameters of at least one first neural network layer in the machine learning model according to the second feature information and the location encoding information of each location point.
[0047] In one possible implementation, the training module is specifically used to update the first parameters based on the expected information and prediction information corresponding to the first data. The first parameters include the weight parameters of the first neural network module and / or the position encoding information of each location point.
[0048] The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the fourth aspect of this application can all be found in the second aspect, and will not be repeated here.
[0049] Fifthly, this application provides an execution device including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; and the processor being used to execute the program in the memory, causing the execution device to perform the method described in the first aspect.
[0050] In a sixth aspect, this application provides a training device including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; and the processor being used to execute the program in the memory, causing the training device to perform the method described in the second aspect above.
[0051] In a seventh aspect, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.
[0052] Eighthly, this application provides a computer program product comprising a program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.
[0053] Ninthly, this application provides a chip system including a processor for supporting the implementation of the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for a terminal device or communication device. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description
[0054] Figure 1 A schematic diagram of the main framework of artificial intelligence provided in the embodiments of this application;
[0055] Figure 2 A system architecture diagram of the data processing system provided in the embodiments of this application;
[0056] Figure 3 A flowchart illustrating a data processing method provided in this application;
[0057] Figure 4 Another flowchart illustrating the data processing method provided in this application;
[0058] Figure 5 A schematic flowchart illustrating a training method for a model provided in an embodiment of this application;
[0059] Figure 6 A schematic diagram illustrating the beneficial effects provided by this application;
[0060] Figure 7 A schematic diagram of a data processing apparatus provided in an embodiment of this application;
[0061] Figure 8 A schematic diagram of a training device for a model provided in an embodiment of this application;
[0062] Figure 9 A schematic diagram of the structure of the execution device provided in the embodiments of this application;
[0063] Figure 10 A schematic diagram of the structure of a training device provided in an embodiment of this application;
[0064] Figure 11 This is a schematic diagram of a chip structure provided in an embodiment of this application. Detailed Implementation
[0065] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will recognize that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0066] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0067] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information (hereinafter referred to as instruction information) is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.
[0068] First, the overall workflow of the artificial intelligence system is described; please refer to [link / reference]. Figure 1 , Figure 1 This is a schematic diagram of a structural framework for the artificial intelligence (AI) subject provided in this application embodiment. The framework is described below from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it could be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT value chain" reflects the value that AI brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.
[0069] (1) Infrastructure
[0070] The infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically employ hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0071] (2) Data
[0072] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0073] (3) Data processing
[0074] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0075] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.
[0076] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0077] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0078] (4) General ability
[0079] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0080] (5) Smart Products and Industry Applications
[0081] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They encapsulate overall artificial intelligence solutions, productize intelligent information decision-making, and realize practical applications. Their application areas mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, intelligent driving, and smart cities.
[0082] The method provided in this application can be applied to various application scenarios that require processing data (hereinafter referred to as "first data") using machine learning models. For example, the first data is multimodal data, which may include images, and may also include text and / or voice. The following are examples of various application scenarios of the method provided in this application.
[0083] 1. Smart City
[0084] For example, in the field of smart cities, it may be necessary to detect water accumulation in the urban environment so that technology can clean it up. Specifically, multimodal data can be acquired, including images and voice information in the urban environment. The voice information could be "Is there water accumulation in the urban environment?" The aforementioned multimodal data can be processed to achieve the detection of water accumulation in the urban environment.
[0085] For example, it may be necessary to locate street vendors in the urban environment in order to deal with them in a timely manner. Specifically, multimodal data can be acquired, including images and text information of the urban environment. The text information can be "Please detect the street vendors in the image". The aforementioned multimodal data is processed to obtain the location information of illegal street vendors in the urban environment. In the field of smart cities, there are many more scenarios that require processing of multimodal data, such as locating pedestrians who run red lights, detecting the behavior of riding electric bikes without helmets in the urban environment, detecting garbage in the urban environment, etc. This application does not exhaustively list them.
[0086] 2. Intelligent driving
[0087] For example, in the field of intelligent driving, vehicles can provide navigation services. For instance, when a user asks which coffee shops are nearby, multimodal data can be obtained. This multimodal data includes a map of a preset area around the user's current location and text information, such as "Which coffee shops are around my current location?" The aforementioned multimodal data can then be processed to obtain the names and locations of several coffee shops around the user.
[0088] 3. Smart Home
[0089] For example, in the field of smart homes, users can provide multimodal data to a refrigerator. Multimodal data includes an image (A) and text information. The image (A) can be an image of a dish, and the text information can be "What ingredients do I need to buy to make this dish?" After processing the multimodal data, the answer corresponding to the text information can be obtained.
[0090] It should be noted that multimodal data processing may also be required in other application scenarios. The examples provided here are only for the purpose of understanding the various application scenarios of the method provided in this application and are not intended to limit this solution.
[0091] Since most of the machine learning models provided in related technologies are for processing single-modal data, in order to process multimodal data, this application discloses that after obtaining the first data to be processed, the first data includes images, and the first data also includes text and / or speech, that is, "first data" can be understood as "multimodal data". Features can be extracted from the images in the first data to obtain the first feature information of the images. Then, based on the first feature information, the weight parameters of at least one first neural network layer in the machine learning model (hereinafter referred to as "first machine learning model" for convenience) are updated to obtain the updated first machine learning model. The updated first machine learning model includes at least one first neural network layer... The updated weight parameters of each first neural network layer, the input of the updated first machine learning model includes text and / or speech in the first data, or can be understood as the input of the first machine learning model including text and / or speech in the first data; thus, the prediction information corresponding to the first data generated by the updated first machine learning model can be obtained, thereby providing a solution capable of processing multimodal data; in addition, in this application, the weight parameters of at least one first neural network layer in the aforementioned machine learning model are updated using the first feature information of the image. This method does not change the length of the text and / or speech input to the machine learning model, so as to minimize the additional computer resource consumption caused by processing multimodal data.
[0092] Before providing a detailed description of the method provided in this application, please refer to [the relevant documentation / reference]. Figure 2 , Figure 2 A system architecture diagram of the data processing system provided in the embodiments of this application is shown. Figure 2 In the process, the data processing system 200 includes a training device 210, a database 220, an execution device 230, a data storage system 240, and a client device 250. The execution device 230 includes a computing module 231.
[0093] In the training phase of the first machine learning model 201, the database 220 stores a training dataset. The training device 210 generates the first machine learning model 201 and uses the training dataset to iteratively train the first machine learning model 201, resulting in a first machine learning model 201 that has undergone training operations. The first machine learning model 201 can be specifically represented as a neural network or as a non-neural network model, depending on the specific application scenario.
[0094] The first machine learning model 201, trained by the training device 210, can be deployed to the computing module 231 of the execution device 230. For example, the execution device 230 can be a computer, mobile phone, tablet, vehicle, smart home device, virtual reality (VR) device, robot, or other type of device. The execution device 230 can access data, code, etc., from the data storage system 240, and can also store data, instructions, etc., in the data storage system 240. The data storage system 240 can be located within the execution device 230, or it can be an external storage device relative to the execution device 230.
[0095] In some embodiments of this application, please refer to Figure 2 The execution device 230 and the client device 250 can be separate independent devices. The execution device 230 is configured with an input / output (I / O) interface to interact with the client device 250. After determining the first information, the client device sends the first information to the execution device 230 through the I / O interface. After the execution device 230 generates the second information corresponding to the first information through the deep learning model 201 in the computing module 231, it can return the aforementioned second information to the client device through the I / O interface.
[0096] It is worth noting that Figure 2 This is merely a schematic diagram of an architecture for a data processing system provided in this embodiment of the invention. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in some other embodiments of this application, the execution device 230 and the client device may be integrated into the same device, allowing the user to directly interact with the execution device 230. Exemplarily, the execution device 230 may be a module in the host CPU of the client device that performs data processing using a deep learning model. The execution device 230 may also be a graphics processing unit (GPU) or a neural network processor (NPU) in the client device, with the GPU or NPU acting as a coprocessor mounted on the host processor, and tasks assigned by the host processor.
[0097] Based on the above description, the specific implementation process of the training and application phases in the method provided in this application will now be described.
[0098] I. Application Phase
[0099] In the embodiments of this application, the first machine learning model used in the application phase can be understood as a first machine learning model that has already undergone training operations; that is, the first machine learning model used in the application phase can be referred to as the trained first machine learning model. For details, please refer to [link to relevant documentation]. Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The data processing method provided in an embodiment of this application may include:
[0100] 301. Obtain first data, which includes images and also includes text and / or voice.
[0101] In this embodiment of the application, the execution device can acquire the first data that needs to be processed. The first data can also be called "multimodal data". The first data includes images, as well as text or voice. The specific information included in the first data can be determined in combination with the actual application scenario.
[0102] 302. Perform feature extraction on the image to obtain the first feature information of the image.
[0103] In this embodiment of the application, after obtaining the first data, the execution device can first use a feature extraction network to extract features from the image to obtain the first feature information of the image. For example, the first feature information may include C feature maps of the same size, where C is an integer greater than or equal to 1. Further, each feature map can be specifically represented as a matrix, and each feature map includes a feature value of each of the n position points. Thus, the C feature maps may include C feature values of each position point, where n is an integer greater than or equal to 1. In this application, "position point" can also be understood as "feature point".
[0104] For example, the feature extraction network can be a convolutional neural network, a residual neural network, a fully connected neural network, an attention-based neural network, or other types of neural networks, which can be determined based on the actual application scenario.
[0105] 303. Based on the first feature information, update the weight parameters of at least one first neural network layer in the first machine learning model to obtain an updated first machine learning model. The updated first machine learning model includes the updated weight parameters of at least one first neural network layer. The input of the updated first machine learning model includes text and / or speech.
[0106] In this embodiment of the application, after obtaining the first feature information, the execution device can update the weight parameters of at least one first neural network layer in the first machine learning model (hereinafter referred to as the "first machine learning model" for convenience). The input of the first machine learning model includes text and / or speech, that is, the first machine learning model is a model for processing text and / or speech.
[0107] The first machine learning model may include at least one neural network module, and each neural network module may include at least one neural network layer. In one case, the aforementioned at least one neural network module includes at least one first feedforward neural network (FFN) module. The at least one first FFN may include all FFNs in the first machine learning model, or the at least one first FFN may also include some FFNs in the first machine learning model. The specific details can be flexibly determined according to the actual situation. The aforementioned at least one first neural network layer is included in the at least one first FFN, that is, the execution device updates the weight parameters of the first neural network layer in the at least one first FFN.
[0108] Optionally, each first neural network layer is a fully connected layer in the first FFN, and a "fully connected layer" can also be called a "linear layer".
[0109] In this embodiment, since the FFN module generally only includes a fully connected layer, that is, the construction of the FFN module is relatively simple. Choosing to update the weight parameters of the first neural network layer in the FFN is beneficial to reducing the complexity of the process of "updating the weight parameters of at least one first neural network layer" and providing a simpler solution.
[0110] Furthermore, since processing multimodal data requires fusing information from different modalities compared to processing only single-modal data, the fusion process consumes additional computer resources (such as storage and computing resources). Because the number of parameters in the weight parameters of the fully connected layer is large, the size of the updated weight parameters of the fully connected layer is not much different from the original weight parameters. Therefore, the impact of using the first feature information of the image to update the weight parameters of the fully connected layer will be smaller, thereby minimizing the additional computer resource consumption during the information fusion of multimodal data.
[0111] For example, in one implementation, before performing step 303, the execution device may also obtain the position encoding information of each of the n position points; then step 303 may include: the execution device updating the weight parameters of at least one first neural network layer in the first machine learning model according to the first feature information and the position encoding information of each position point.
[0112] In this application, "location encoding information of each location point" means that the location encoding information of the location point is related to the location of the location point. Optionally, the location encoding information of each location point can be learned during the training phase, and the location encoding information of different location points is different; or, the location encoding information of each location point can also be obtained using a preset location encoding algorithm, which can be determined according to the actual application scenario.
[0113] The location encoding information for each location point may include first location encoding information and second location encoding information for that location point. Optionally, both the first and second location encoding information for each location point can be learned during the training phase, in which case the first location encoding information and the second location encoding information for different location points can be different. Optionally, the first location encoding information and the second location encoding information for each location point can also be different.
[0114] Alternatively, the location encoding information for each location point may include only one location encoding information (hereinafter referred to as "third location encoding information" for easy distinction); optionally, the third location encoding information for each location point may be learned during the training phase, in which case the third location encoding information for different location points will be different.
[0115] For example, the first position encoding information (or the second position encoding information, or the third position encoding information) of each position point can be specifically represented as a first vector.
[0116] After acquiring the first feature information of the image, the execution device can also transform the first feature information into second feature information through the first neural network module, and then update the weight parameters of at least one first neural network layer in the first machine learning model based on the second feature information and the position encoding information of each location point.
[0117] The second feature information can be in the form of a token, meaning the form of the second feature information matches that of the text or speech feature information. For example, the second feature information may include first sub-feature information for each location point. Specifically, the first sub-feature information for each location point can be represented as a second vector. For instance, the second vector may include C feature values for each location point. Optionally, the first vector and the second vector have equal lengths; that is, the first vector may also include C values.
[0118] For example, the first neural network module may include at least one fully connected layer, which may also be called a "linear layer"; and / or, the first neural network module may also include a multilayer perceptron; and / or, the first neural network module may also include a convolutional layer; and / or, the first neural network module may also be a neural network layer based on an attention mechanism, etc., and the specifics can be determined in combination with the actual application scenario.
[0119] Optionally, in order to minimize the computer resources consumed when processing multimodal data, the number of parameters of the first neural network module can be within a first preset range, which can be less than or equal to 10 million parameters.
[0120] Regarding the specific implementation of the execution device "updating the weight parameters of at least one first neural network layer in the first machine learning model based on the second feature information and the position encoding information of each location point", in one case, the execution device may update the weight parameters of the first fully connected layer and the second fully connected layer in each first FFN module included in the first machine learning model.
[0121] Specifically, in one implementation, the location encoding information of each location point may include first location encoding information and second location encoding information of each location point. The execution device can fuse the first sub-feature information and the first location encoding information of each location point to obtain the second sub-feature information of each location point, and fuse the first sub-feature information and the second location encoding information of each location point to obtain the third sub-feature information of each location point. Then, the execution device uses the second sub-feature information of each location point to update the weight parameters of the first fully connected layer to obtain the updated weight parameters of the first fully connected layer, and uses the third sub-feature information of each location point to update the weight parameters of the second fully connected layer to obtain the updated weight parameters of the second fully connected layer. The updated weight parameters of the first FFN include the updated weight parameters of the first fully connected layer and the updated weight parameters of the second fully connected layer. The execution device performs the aforementioned operations on each of the at least one first FFN module to obtain the updated weight parameters of each first FFN.
[0122] Optionally, the execution device may fuse the first sub-feature information and the first position encoding information of each location point, which may include: the execution device adding the first sub-feature information and the first position encoding information of each location point; or the execution device may fuse the first sub-feature information and the second position encoding information of each location point, which may include: the execution device adding the first sub-feature information and the second position encoding information of each location point. That is, the above fusion operation can be addition. Alternatively, the aforementioned "addition" can be replaced by "multiplication", "subtraction" or other fusion operations, etc.
[0123] In this application embodiment, a specific implementation scheme is provided for fusing the first sub-feature information of each location point and the first location encoding information (or the second location encoding information) of each location point. This improves the feasibility of the scheme and the implementation scheme is relatively simple, thereby reducing the difficulty of fusing data between different modalities and helping to reduce the additional computer resource consumption caused by fusing data between different modalities.
[0124] Furthermore, in implementation method (1), the execution device updates the weight parameters of the first fully connected layer using the second sub-feature information of each location point, which may include: the execution device concatenating the weight parameters of the first fully connected layer with the second sub-feature information of each location point; the execution device updates the weight parameters of the second fully connected layer using the third sub-feature information of each location point, which may include: the execution device concatenating the weight parameters of the second fully connected layer with the third sub-feature information of each location point.
[0125] A first FFN may include at least two fully connected layers. To further understand this scheme, two examples of the first FFN are provided below. In Example 1, the first FFN can be composed of two fully connected layers, and the first FFN can be expressed as:
[0126] FFN(x)=φ(xW1)W2
[0127] W1 represents the weight parameters of the first fully connected layer, and W2 represents the weight parameters of the second fully connected layer. The specific expressions for W1 and W2 can be as follows:
[0128] W1 = (k1, k2, ..., k D W2 = (v1, v2, ..., v) D ) T
[0129] The second and third sub-feature information for each location point can be as follows:
[0130]
[0131]
[0132] Among them, z j K(z) represents any one of the n possible locations. j ) represents z j The second sub-feature information, f(z) j ) represents z j The first sub-feature information, where λ is a hyperparameter. Represents z j The first position encoding information, V(z) j ) represents z j The third sub-feature information, Represents z j The second position encoding information.
[0133] The updated weight matrices of the first fully connected layer and the second fully connected layer can be as follows:
[0134]
[0135]
[0136] Where W′1 represents the updated weight parameters of the first fully connected layer, and W′2 represents the updated weight parameters of the second fully connected layer.
[0137] Referring to the above formula, the updated weight parameters of the first fully connected layer are obtained by concatenating W1 (i.e., the weight parameters of the first fully connected layer) with the second sub-feature information of each of the n position points; the updated weight parameters of the first fully connected layer are obtained by concatenating W2 (i.e., the weight parameters of the second fully connected layer) with the second sub-feature information of each of the n position points.
[0138] For example, the weight parameters of the first fully connected layer can be represented in matrix form, and the updated weight parameters of the first fully connected layer can also be represented in matrix form; correspondingly, the weight parameters of the second fully connected layer can be represented in matrix form, and the updated weight parameters of the second fully connected layer can also be represented in matrix form.
[0139] Optionally, the updated weight parameters of the first fully connected layer can be A multiplied by B, and the updated weight parameters of the second fully connected layer can be B multiplied by A, where A and B are both integers greater than or equal to 1. That is, the size of W′1 can be the same as the size of the transpose of W′2, so that the size of the information generated by the first FFN is the same as the size of the information generated by the updated first FFN, thereby further reducing the interference caused by processing multimodal data and ensuring the stability of the first machine learning model operation process.
[0140] In Example 2, the first FFN can include three fully connected layers, and the first FFN can be expressed as:
[0141]
[0142] The specific expressions for W1 and W2 can be found in the formulas above, and will not be repeated here. The specific expression for W3 is as follows:
[0143] W3 = (g1, ..., g D ),
[0144] The first fully connected layer can be either W1 or W3, and the second fully connected layer can be W2.
[0145] Optionally, the size of the updated weight parameters of the first fully connected layer can be A multiplied by B, and the size of the updated weight parameters of the second fully connected layer can be B multiplied by A. In addition, since there is a third fully connected layer in Example 2, in order to ensure that the size of the information generated by the first FFN is the same as the size of the information generated by the updated first FFN, the weight parameters of the third fully connected layer can also be updated so that x multiplied by the updated weight parameters of the third fully connected layer is always equal to 1. It should be understood that Examples 1 and 2 listed here are only for the convenience of understanding this scheme and are not intended to limit this scheme.
[0146] In this embodiment, the weight parameters of the first fully connected layer are updated by concatenating the weight parameters of the first fully connected layer with the second sub-feature information of each location point, and the weight parameters of the second fully connected layer are updated by concatenating the weight parameters of the second fully connected layer with the third sub-feature information of each location point. This preserves the weight parameters of the first and second fully connected layers before the aforementioned update, and also incorporates information obtained from the image into the updated weight parameters of the first and second fully connected layers. Since the weight parameters of the first and second fully connected layers were trained before the aforementioned update, the accuracy of the prediction information obtained when processing text and / or speech using the weight parameters of the first and second fully connected layers before the update is high. When updating the weight parameters of the first and second fully connected layers using information obtained from the image, the weight parameters of the first and second fully connected layers before the update are preserved, which helps to ensure that the prediction information generated by the updated first machine learning model also maintains high accuracy.
[0147] In implementation method (2), the weight parameters of the first fully connected layer can be represented as a matrix of the first size, and the weight parameters of the second fully connected layer can be represented as a matrix of the second size. The execution device updates the weight parameters of the first fully connected layer using the second sub-feature information of each location point. This can include: after obtaining the second sub-feature information of each of the n location points, the execution device can place the second sub-feature information of each of the n location points into a first matrix of the first size. The empty values in the first matrix can be filled with 0. Then, the weight parameters of the first fully connected layer and the first matrix are added together to obtain the updated weight parameters of the first fully connected layer. Correspondingly, the execution device updates the weight parameters of the second fully connected layer using the third sub-feature information of each location point. This can include: after obtaining the third sub-feature information of each of the n location points, the execution device can place the third sub-feature information of each of the n location points into a second matrix of the second size. The empty values in the second matrix can be filled with 0. Then, the weight parameters of the second fully connected layer and the second matrix are added together to obtain the updated weight parameters of the second fully connected layer.
[0148] For example, the above "addition" can also be replaced by "subtraction" or other operations, as long as the size of the weight parameters of the first fully connected layer (or the second fully connected layer) remains unchanged before and after the update.
[0149] In this embodiment, since the FFN module generally has a relatively simple structure, the position encoding information of each position point includes first position encoding information and second position encoding information. The first sub-feature information of each position point is fused with the first position encoding information and the second position encoding information to obtain the second sub-feature information and the third sub-feature information of each position point. Then, the weight parameters of the first fully connected layer and the second fully connected layer are updated using the second sub-feature information and the third sub-feature information of each position point. Since updating the weight parameters of the first neural network layer may cause the size of the weight parameters of the first neural network layer to change, which in turn causes the size of the information generated by the first neural network layer to change, choosing to update the weight parameters of the first fully connected layer and the second fully connected layer in the FFN is beneficial to achieve mutual cancellation between the influence of the first fully connected layer on the size of the processed information and the influence of the second fully connected layer on the size of the processed information. This is beneficial to ensure that the size of the information generated by the FFN module remains unchanged, thereby ensuring the smoothness of the updated first machine learning model in the data processing process.
[0150] In another implementation, the positional encoding information of each location point may include the third positional encoding information of each location point. The execution device can then fuse the first sub-feature information and the third positional encoding information of each location point to obtain the fourth sub-feature information of each location point. Furthermore, the execution device uses the fourth sub-feature information of each location point to update the weight parameters of the first fully connected layer, obtaining the updated weight parameters of the first fully connected layer. It also uses the third positional encoding information of each location point to update the weight parameters of the second fully connected layer, obtaining the updated weight parameters of the second fully connected layer. The updated weight parameters of the first FFN include the updated weight parameters of the first and second fully connected layers. The execution device performs the aforementioned operations on each of the at least one first FFN module to obtain the updated weight parameters of each first FFN.
[0151] For example, the above fusion operation can be addition, multiplication, subtraction or other operations, which are not limited here.
[0152] For example, the specific implementation of "the execution device updates the weight parameters of the first fully connected layer using the fourth sub-feature information of each location point" can be referred to the above description of "the execution device updates the weight parameters of the first fully connected layer using the second sub-feature information of each location point", the difference being that "second sub-feature information" in the above description is replaced with "fourth sub-feature information"; the specific implementation of "the execution device updates the weight parameters of the second fully connected layer using the third position encoding information of each location point" can be referred to the above description of the specific implementation of "the execution device updates the weight parameters of the first fully connected layer using the third sub-feature information of each location point", the difference being that "third sub-feature information" in the above description is replaced with "third position encoding information", which will not be elaborated here.
[0153] In one scenario, at least one first neural network layer can be a neural network layer randomly selected from the first machine learning model, or at least one first neural network layer can be a neural network layer of a predetermined number of layers in the first machine learning model, such as the 20th, 30th, and 40th layers. This example is provided for ease of understanding. The execution device can fuse the first sub-feature information and the third positional encoding information of each position point to obtain the fourth sub-feature information of each position point. Then, it uses the fourth sub-feature information of each position point to update the weight parameters of each first neural network layer, obtaining the updated weight parameters of each first neural network layer, thereby obtaining the weight parameters of each of the at least one first neural network layers.
[0154] For example, the weight parameters of a first neural network layer can be represented as a matrix of third size. The execution device can put the fourth sub-feature information of each of the n location points into the third matrix of third size, and pad the third matrix with zeros. The weight parameters of the first neural network layer are added to the third matrix to obtain the updated weight parameters of the first neural network layer. For example, the aforementioned "addition" can also be replaced by "subtraction" or other operations, as long as the size of the weight parameters of the first neural network layer and the updated weight parameters of the first neural network layer are consistent.
[0155] In this embodiment of the application, when updating the weight parameters of at least one first neural network layer, not only is the first feature information in the form of feature map used, but also the position encoding information of each location point is obtained. This allows for the use of more information obtained from the image to jointly update the weight parameters of at least one first neural network layer, which is beneficial for making fuller use of the information obtained from the image and thus for obtaining prediction information that is more compatible with the first data, thereby improving the accuracy of the obtained prediction information.
[0156] In another implementation, the execution device may not acquire the location encoding information of each location point, but instead convert the first feature information into second feature information through the first neural network module. The second feature information is feature information in the form of a token and includes the first sub-feature information of each location point. Then, the execution device may update the weight parameters of at least one first neural network layer in the machine learning model using only the second feature information.
[0157] For example, in one case, the first machine learning model includes at least one first FFN, and at least one first neural network layer includes a first fully connected layer and a second fully connected layer in each first FFN. Step 303 may include: the execution device updating the weight parameters of the first fully connected layer using the first sub-feature information of each location point to obtain the updated weight parameters of the first fully connected layer, and updating the weight parameters of the second fully connected layer using the first sub-feature information of each location point to obtain the updated weight parameters of the second fully connected layer.
[0158] The first machine learning model includes at least one first FFN module, each of the at least one first FFN module includes a first fully connected layer and a second fully connected layer, the updated weight parameters of at least one first neural network layer include the updated weight parameters of each first FFN, and the updated weight parameters of each first FFN include the updated weight parameters of the first fully connected layer and the updated weight parameters of the second fully connected layer.
[0159] For example, the specific implementation of "the execution device updates the weight parameters of the first fully connected layer using the first sub-feature information of each location point" can be referred to the above description of "the execution device updates the weight parameters of the first fully connected layer using the second sub-feature information of each location point", the difference being that "second sub-feature information" in the above description is replaced with "first sub-feature information"; the specific implementation of "the execution device updates the weight parameters of the second fully connected layer using the first sub-feature information of each location point" can be referred to the above description of "the execution device updates the weight parameters of the first fully connected layer using the third sub-feature information of each location point", the difference being that "third sub-feature information" in the above description is replaced with "first sub-feature information", and will not be repeated here.
[0160] In another scenario, at least one first neural network layer may be a neural network layer randomly selected from the first machine learning model, or at least one first neural network layer may be a neural network layer with a predetermined number of layers in the first machine learning model. Step 303 may include: the execution device updating the weight parameters of each first neural network layer using the first sub-feature information of each location point, to obtain the updated weight parameters of each first neural network layer; the specific implementation of the execution device performing the aforementioned steps can be found in the above description of the specific implementation of "the execution device updating the weight parameters of each first neural network layer using the fourth sub-feature information of each location point", the difference being that "fourth sub-feature information" in the above description is replaced with "first sub-feature information", which will not be elaborated here.
[0161] Since the first feature information is in the form of a feature map, and in this application, the information obtained from the image is integrated into the processing of the first machine learning model used to process text and / or speech, the first feature information is converted into second feature information in the form of a token. The second feature information in the form of a token is consistent with the form of the feature information of text and / or speech. Then, the weight parameters of at least one first neural network layer are updated using the second feature information in the form of a token and the position encoding information of each location point. This helps to reduce the difficulty of the process of "updating the weight parameters of at least one first neural network layer", improve the degree of fusion between multimodal data, and help to obtain prediction information that is more suitable for the first data, that is, to improve the accuracy of the obtained prediction information.
[0162] 304. Obtain the prediction information corresponding to the first data. The prediction information is generated by the updated first machine learning model.
[0163] In this embodiment, after updating the weight parameters of at least one first neural network layer in the first machine learning model, the execution device can also obtain prediction information corresponding to the first data generated by the updated first machine learning model; wherein, the input of the first machine learning model includes text and / or speech in the first data, that is, the input of the updated first machine learning model includes text and / or speech in the first data. For example, the prediction information corresponding to the first data can be text information or speech information.
[0164] For a more intuitive understanding of this solution, please refer to [link / reference]. Figure 4 , Figure 4 Another flowchart illustrating the data processing method provided in this application, as shown below. Figure 4As shown, the execution device acquires first data including images and text. The execution device inputs the image from the first data into a feature extraction network to obtain first feature information generated by the feature extraction network. Then, it uses a first neural network module to transform the first feature information to obtain second feature information. The second feature information is represented as token-based feature information, which includes first sub-feature information for each location point. The execution device adds the first sub-feature information of each location point to the first position encoding information of each location point to obtain second sub-feature information for each location point. Then, it uses the second sub-feature information of each location point to update the weight parameters of the first fully connected layer in the FFN to obtain the updated weight parameters of the first fully connected layer. The execution device adds the first sub-feature information of each location point to the second position encoding information of each location point to obtain third sub-feature information for each location point. Then, it uses the third sub-feature information of each location point to update the weight parameters of the second fully connected layer in the FFN to obtain the updated weight parameters of the second fully connected layer.
[0165] The execution device can vectorize the text in the first data to obtain the initial feature information of the text, which can also be understood as text in token form. The execution device inputs the text in token form into the first machine learning model, and after updating the first machine learning model, obtains the prediction information generated by the updated first machine learning model. Figure 4 Taking the first machine learning model in China, which includes a neural network module based on multi-head self-attention (MHSA) and an FFN module, as an example, it should be understood that... Figure 4 The examples in the text are only for the purpose of understanding this solution. The aforementioned MHSA can also be replaced with a recurrent neural network module, a residual neural network module, or other types of neural network modules, etc. The specific method can be determined based on the actual application scenario.
[0166] In this embodiment, after acquiring the first multimodal data, which includes images and text / / or speech, feature extraction can be performed on the images to obtain first feature information. This first feature information is then used to update the weight parameters of at least one first neural network layer in a machine learning model, resulting in an updated first machine learning model. The input to this first machine learning model includes text and / or speech. The updated first machine learning model generates prediction information corresponding to the first data. Therefore, in generating the prediction information, not only image information but also text and / or speech information is used, providing a solution capable of processing multimodal data. Furthermore, this application uses the first feature information of the images to update the weight parameters of at least one first neural network layer in the aforementioned machine learning model. This method does not change the length of the text and / or speech input to the machine learning model, minimizing the additional computer resource consumption caused by processing multimodal data.
[0167] II. Training Phase
[0168] Optionally, in this embodiment of the application, the first machine learning model can be trained first, and after obtaining the first machine learning model that has undergone training operations, the following can be performed: Figure 5 The illustrated embodiment; alternatively, the feature extraction network used for image feature extraction can be trained first, and after obtaining the trained feature extraction network, the following can be performed. Figure 5 For the specific embodiments shown, please refer to [link / reference]. Figure 5 , Figure 5 This is a flowchart illustrating a training method for a model provided in an embodiment of this application. The training method for a model provided in an embodiment of this application may include:
[0169] 501. Obtain first data, which includes images and also includes text and / or voice.
[0170] In this embodiment of the application, during the training phase, a training data set may be deployed in the training device. The training data set includes multiple first data sets and expected information corresponding to each first data set. The "expected information corresponding to the first data set" represents the correct information corresponding to the first data set. The "expected information corresponding to the first data set" may also be called the "label corresponding to the first data set" or the "truth value corresponding to the first data set", etc.
[0171] 502. Perform feature extraction on the image to obtain the first feature information of the image.
[0172] 503. Based on the first feature information, update the weight parameters of at least one first neural network layer in the first machine learning model to obtain an updated first machine learning model. The updated first machine learning model includes the updated weight parameters of at least one first neural network layer. The input of the updated first machine learning model includes text and / or speech.
[0173] 504. Obtain the prediction information corresponding to the first data. The prediction information is generated by the updated first machine learning model.
[0174] In this embodiment, the specific implementation of steps 502 to 504 of the training device can be found in the above description. Figure 3 The descriptions of steps 302 to 304 in the corresponding embodiments will not be repeated here.
[0175] 505. Training is performed based on the expected information and predicted information corresponding to the first data. The training objectives include improving the similarity between the predicted information and the expected information.
[0176] In this embodiment of the application, for example, the "expected information corresponding to the first data" and the "prediction information" can be the same type of data. For example, if the prediction information is text information, then the "expected information corresponding to the first data" is text information; if the prediction information is voice information, then the "expected information corresponding to the first data" is voice information.
[0177] For example, after generating the prediction information corresponding to the first data in step 504, the training device can generate the function value of the loss function based on the expected information and the prediction information corresponding to the first data. Based on the function value of the loss function, the first parameter is updated using the backpropagation algorithm to complete one training cycle. The loss function can represent the similarity between the expected information and the prediction information corresponding to the first data. The goal of training using the loss function can include improving the similarity between the prediction information and the expected information.
[0178] The training device repeatedly executes steps 501 to 505 multiple times to update the first parameter multiple times (i.e., to perform multiple training sessions) until the first condition is met, at which point the first parameter that has undergone training is obtained; the first condition may be the convergence condition of the loss function, or the training here reaching a preset number of times, etc.
[0179] Optionally, the first parameter includes the weight parameters of the first neural network module and / or the position encoding information of each location point. The meanings of "first neural network module" and "position encoding information of each location point" can be found above. Figure 3The descriptions in the corresponding embodiments will not be repeated here. Since the first machine learning model is a model for processing text and / or speech, in order to provide a solution for processing multimodal data, the feature information of the image is integrated into the processing of the first machine learning model in this application. Optionally, a first neural network module and the position encoding information of each location point are added in the aforementioned fusion process. During the training phase, the weight parameters of the first neural network module and / or the position encoding information of each location point are adjusted, thereby reducing the number of parameters that need to be adjusted during the training phase. This not only helps to reduce the storage resources consumed during the training phase, but also helps to improve the efficiency of the training phase.
[0180] For example, the first parameter may also include weight parameters of other neural network layers in the first machine learning model. For instance, the first parameter may also include weight parameters of the normalization layer in the first machine learning model. The specific parameters included in the first parameter can be determined in conjunction with the actual application scenario.
[0181] To gain a more intuitive understanding of the beneficial effects of the method provided in this application, please refer to [link / reference needed]. Figure 6 , Figure 6 A schematic diagram illustrating the beneficial effects provided by this application. For example... Figure 6 As shown, the feature extraction network in ResNet101 is used to extract features from images. That is, ResNet101 represents a model for image processing. BART-base and T5-base are used as machine learning models to process text information. VQAv2 and GQA represent two different datasets. VQA Score represents an indicator of the accuracy of the final prediction information. The higher the VQA Score, the better the accuracy of the prediction information. FLOPs(G) represents the computer resources consumed when processing multimodal data. The smaller the FLOPs(G), the less computer resources are consumed.
[0182] Here, Compacter, LoRA, and VL-PET represent three schemes for processing multimodal data in related technologies, and MemVP represents processing multimodal data using the method provided in this application. See [link to relevant documentation]. Figure 6 It is evident that the method provided in this application requires the fewest parameters to be adjusted during the training phase, achieves a high VQA score, and consumes fewer computer resources.
[0183] exist Figures 1 to 6 Based on the corresponding embodiments, in order to better implement the above-described solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 7 , Figure 7This is a schematic diagram of a data processing apparatus provided in an embodiment of this application. The data processing apparatus 700 includes: an acquisition module 701, used to acquire first data, the first data including an image, and the first data further including text and / or speech; a feature extraction module 702, used to extract features from the image to obtain first feature information of the image; an update module 703, used to update the weight parameters of at least one first neural network layer in a machine learning model according to the first feature information to obtain an updated machine learning model, the updated machine learning model including the updated weight parameters of at least one first neural network layer, and the input of the updated machine learning model including text and / or speech; the acquisition module 701 is also used to acquire prediction information corresponding to the first data, the prediction information being generated by the updated machine learning model.
[0184] Optionally, the first feature information is a feature map, and there is at least one location point in the feature map. The acquisition module 701 is also used to acquire the location encoding information of each location point in the at least one location point. The update module 703 is specifically used to update the weight parameters of at least one first neural network layer in the machine learning model according to the first feature information and the location encoding information of each location point.
[0185] Optionally, at least one first neural network layer is a neural network layer in the first feedforward neural network (FFN) module included in the machine learning model.
[0186] Optionally, the first neural network layer is a fully connected layer in the first FFN module included in the machine learning model.
[0187] Optionally, the update module 703 is specifically used to: convert the first feature information into second feature information through the first neural network module, wherein the second feature information is feature information in the form of a token; and update the weight parameters of at least one first neural network layer in the machine learning model according to the second feature information and the position encoding information of each location point.
[0188] Optionally, the second feature information includes first sub-feature information for each location point, and the location encoding information for each location point includes first location encoding information and second location encoding information. The update module 703 is specifically used to: fuse the first sub-feature information and the first location encoding information for each location point to obtain the second sub-feature information for each location point; fuse the first sub-feature information and the second location encoding information for each location point to obtain the third sub-feature information for each location point; update the weight parameters of the first fully connected layer using the second sub-feature information for each location point to obtain the updated weight parameters of the first fully connected layer; and update the weight parameters of the second fully connected layer using the third sub-feature information for each location point to obtain the updated weight parameters of the second fully connected layer. The machine learning model includes at least one first FFN module, the first FFN module includes a first fully connected layer and a second fully connected layer, and the updated weight parameters of at least one first neural network layer include the updated weight parameters of each first FFN.
[0189] Optionally, the update module 703 is specifically used to: concatenate the weight parameters of the first fully connected layer with the second sub-feature information of each location point; and concatenate the weight parameters of the second fully connected layer with the third sub-feature information of each location point.
[0190] Optionally, the update module 703 is specifically used to: add the first sub-feature information of each location point to the first location encoding information of each location point; and add the second sub-feature information of each location point to the second location encoding information of each location point.
[0191] Optionally, during the training phase, the parameters that need to be adjusted include: the weight parameters of the first neural network module and / or the location encoding information of each location point.
[0192] It should be noted that the information interaction and execution process between the modules / units in the data processing device 700 are different from those in this application. Figures 1 to 6 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0193] See Figure 8 , Figure 8This is a schematic diagram of a model training device provided in an embodiment of this application. The model training device 800 includes: an acquisition module 801, used to acquire first data, the first data including an image, and the first data further including text and / or speech; a feature extraction module 802, used to extract features from the image to obtain first feature information of the image; an update module 803, used to update the weight parameters of at least one first neural network layer in the machine learning model according to the first feature information to obtain an updated machine learning model, the updated machine learning model including the updated weight parameters of at least one first neural network layer, and the input of the updated machine learning model including text and / or speech; the acquisition module 801 is also used to acquire prediction information corresponding to the first data, the prediction information being generated by the updated machine learning model; and a training module 804, used to train according to the expected information and prediction information corresponding to the first data, the training objective including improving the similarity between the prediction information and the expected information.
[0194] Optionally, the first feature information is a feature map, in which at least one location point exists. The acquisition module 801 is further used to acquire the location encoding information of each location point in the at least one location point. The update module is specifically used to update the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the location encoding information of each location point.
[0195] Optionally, the update module 803 is specifically used to: convert the first feature information into second feature information through the first neural network module, wherein the second feature information is feature information in the form of a token; and update the weight parameters of at least one first neural network layer in the machine learning model according to the second feature information and the position encoding information of each position point.
[0196] Optionally, the training module 804 is specifically used to update the first parameters based on the expected information and prediction information corresponding to the first data. The first parameters include the weight parameters of the first neural network module and / or the position encoding information of each location point.
[0197] It should be noted that the information interaction and execution process between the modules / units in the model training device 800 are different from those in this application. Figures 1 to 6 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0198] The following describes an execution device provided in an embodiment of this application. Please refer to [link / reference]. Figure 9 , Figure 9This is a schematic diagram of an execution device provided in an embodiment of this application. Specifically, the execution device 900 includes: a receiver 901, a transmitter 902, a processor 903, and a memory 904 (wherein the execution device 900 may have one or more processors 903). Figure 9 (Taking a processor as an example), processor 903 may include application processor 9031 and communication processor 9032. In some embodiments of this application, receiver 901, transmitter 902, processor 903 and memory 904 may be connected via a bus or other means.
[0199] Memory 904 may include read-only memory and random access memory, and provides instructions and data to processor 903. A portion of memory 904 may also include non-volatile random access memory (NVRAM). Memory 904 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0200] Processor 903 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.
[0201] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 903. The processor 903 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 903 or by instructions in software form. The processor 903 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 903 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 904, and processor 903 reads the information from memory 904 and, in conjunction with its hardware, completes the steps of the above method.
[0202] Receiver 901 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 902 can be used to output digital or character information through the first interface; transmitter 902 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 902 may also include a display device such as a display screen.
[0203] In this embodiment, the processor 903 is used to execute... Figures 1 to 6 The method executed by the execution device in the corresponding embodiment. It should be noted that the specific manner in which the application processor 9031 in processor 903 executes the aforementioned steps differs from that in this application. Figures 1 to 6 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 6 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0204] This application also provides a training device; please refer to [link / reference]. Figure 10 , Figure 10This is a schematic diagram of a training device provided in an embodiment of this application. Specifically, the training device 1000 is implemented by one or more servers. The training device 1000 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1022 (e.g., one or more processors) and a memory 1032, and one or more storage media 1030 (e.g., one or more mass storage devices) for storing application programs 1042 or data 1044. The memory 1032 and storage media 1030 can be temporary or persistent storage. The program stored in the storage media 1030 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the training device. Furthermore, the CPU 1022 may be configured to communicate with the storage media 1030 and execute the series of instruction operations in the storage media 1030 on the training device 1000.
[0205] The training device 1000 may also include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0206] In this embodiment, the central processing unit 1022 is used to execute... Figures 1 to 6 The method executed by the training device in the corresponding embodiment. It should be noted that the specific manner in which the central processing unit 1022 executes the above steps differs from that in this application. Figures 1 to 6 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 6 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0207] This application embodiment also provides a computer-readable storage medium storing a program for performing signal processing, which, when run on a computer, causes the computer to perform the aforementioned actions. Figures 1 to 6 The method described in the illustrated embodiment executes steps performed by the device, or causes the computer to perform steps as described above. Figures 1 to 6 The steps performed by the training device in the method described in the illustrated embodiment.
[0208] This application also provides a computer program product, which includes a program that, when run on a computer, causes the computer to perform the aforementioned actions. Figures 1 to 6 The method described in the illustrated embodiment executes steps performed by the device, or causes the computer to perform steps as described above. Figures 1 to 6 The steps performed by the training device in the method described in the illustrated embodiment.
[0209] The execution device and training device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip to perform the above-mentioned operations. Figures 1 to 6 Method 1 is described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0210] For details, please refer to Figure 11 , Figure 11 This is a schematic diagram of a chip provided in an embodiment of this application. The chip can be represented as a neural network processor (NPU) 110. The NPU 110 is mounted as a coprocessor on the host CPU, and tasks are assigned by the host CPU. The core part of the NPU is the arithmetic circuit 1103, which is controlled by the controller 1104 to extract matrix data from the memory and perform multiplication operations.
[0211] In some implementations, the arithmetic circuit 1103 internally includes multiple processing engines (PEs). In some implementations, the arithmetic circuit 1103 is a two-dimensional pulsating array. The arithmetic circuit 1103 can also be a one-dimensional pulsating array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1103 is a general-purpose matrix processor.
[0212] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 1102 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 1101 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 1108.
[0213] Unified memory 1106 is used to store input and output data. Weight data is directly transferred to weight memory 1102 via Direct Memory Access Controller (DMAC) 1105. Input data is also transferred to unified memory 1106 via DMAC.
[0214] BIU stands for Bus Interface Unit, which is used for interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1109.
[0215] The Bus Interface Unit (BIU) 1110 is used by the instruction fetch memory 1109 to fetch instructions from external memory, and also by the memory access controller 1105 to fetch the original data of the input matrix A or the weight matrix B from external memory.
[0216] The DMAC is mainly used to move input data from external memory DDR to unified memory 1106, or to move weight data to weight memory 1102, or to move input data to input memory 1101.
[0217] The vector computation unit 1107 includes multiple arithmetic processing units that, when needed, further process the output of the computation circuit, such as vector multiplication, vector addition, exponential operations, logarithmic operations, size comparisons, etc. It is mainly used for computation in non-convolutional / fully connected layers of neural networks, such as Batch Normalization, pixel-level summation, and upsampling of feature planes.
[0218] In some implementations, the vector computation unit 1107 can store the processed output vector in the unified memory 1106. For example, the vector computation unit 1107 can apply linear and / or nonlinear functions to the output of the computation circuit 1103, such as linear interpolation of feature planes extracted by convolutional layers, or, for example, accumulating a vector of values to generate activation values. In some implementations, the vector computation unit 1107 generates normalized values, pixel-level summed values, or both. In some implementations, the processed output vector can be used as activation input to the computation circuit 1103, for example, for use in subsequent layers of the neural network.
[0219] The instruction fetch buffer 1109 connected to the controller 1104 is used to store the instructions used by the controller 1104;
[0220] Unified memory 1106, input memory 1101, weight memory 1102, and instruction fetch memory 1109 are all on-chip memories. External memory is proprietary to this NPU hardware architecture.
[0221] The machine learning model and the operations of each layer in the neural network shown in the above embodiments can be executed by the operation circuit 1103 or the vector calculation unit 1107.
[0222] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.
[0223] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0224] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0225] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0226] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A data processing method, characterized in that, The method includes: Acquire first data, the first data including images, and the first data further including text and / or voice; Feature extraction is performed on the image to obtain the first feature information of the image; Based on the first feature information, the weight parameters of at least one first neural network layer in the machine learning model are updated to obtain an updated machine learning model. The updated machine learning model includes the updated weight parameters of the at least one first neural network layer, and the input of the updated machine learning model includes the text and / or speech. Obtain prediction information corresponding to the first data, wherein the prediction information is generated by the updated machine learning model.
2. The method according to claim 1, characterized in that, The first feature information is a feature map, and the feature map contains at least one location point. The method further includes: obtaining the location encoding information of each location point among the at least one location point. The step of updating the weight parameters of at least one first neural network layer in the machine learning model according to the first feature information includes: updating the weight parameters of the at least one first neural network layer in the machine learning model according to the first feature information and the position encoding information of each position point.
3. The method according to claim 1 or 2, characterized in that, The at least one first neural network layer is a neural network layer in the first feedforward neural network (FFN) module included in the machine learning model.
4. The method according to claim 3, characterized in that, The first neural network layer is a fully connected layer in the first FFN module included in the machine learning model.
5. The method according to claim 2, characterized in that, The step of updating the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the position encoding information of each location point includes: The first feature information is converted into second feature information by the first neural network module, and the second feature information is feature information in the form of a token. The weight parameters of at least one first neural network layer in the machine learning model are updated based on the second feature information and the location encoding information of each location point.
6. The method according to claim 5, characterized in that, The second feature information includes first sub-feature information for each location point, and the location encoding information for each location point includes first location encoding information and second location encoding information. Updating the weight parameters of at least one first neural network layer in the machine learning model based on the second feature information and the location encoding information of each location point includes: The first sub-feature information of each location point and the first location encoding information of each location point are fused to obtain the second sub-feature information of each location point; The first sub-feature information and the second position encoding information of each position point are fused to obtain the third sub-feature information of each position point; The weight parameters of the first fully connected layer are updated using the second sub-feature information of each location point to obtain the updated weight parameters of the first fully connected layer. The weight parameters of the second fully connected layer are updated using the third sub-feature information of each location point to obtain the updated weight parameters of the second fully connected layer. The machine learning model includes at least one first FFN module, the first FFN module includes the first fully connected layer and the second fully connected layer, and the updated weight parameters of the at least one first neural network layer include the updated weight parameters of each first FFN.
7. The method according to claim 6, characterized in that, The step of updating the weight parameters of the first fully connected layer using the second sub-feature information of each location point includes: concatenating the weight parameters of the first fully connected layer with the second sub-feature information of each location point; The step of updating the weight parameters of the second fully connected layer using the third sub-feature information of each location point includes: concatenating the weight parameters of the second fully connected layer with the third sub-feature information of each location point.
8. The method according to claim 6 or 7, characterized in that, The step of fusing the first sub-feature information of each location point and the first location encoding information of each location point includes: adding the first sub-feature information of each location point and the first location encoding information of each location point; The step of fusing the first sub-feature information and the second position encoding information of each position point includes: adding the second sub-feature information and the second position encoding information of each position point.
9. The method according to any one of claims 5 to 7, characterized in that, During the training phase, the parameters that need to be adjusted include: the weight parameters of the first neural network module and / or the position encoding information of each location point.
10. A method for training a model, characterized in that, The method includes: Acquire first data, the first data including images, and the first data further including text and / or voice; Feature extraction is performed on the image to obtain the first feature information of the image; Based on the first feature information, the weight parameters of at least one first neural network layer in the machine learning model are updated to obtain an updated machine learning model. The updated machine learning model includes the updated weight parameters of the at least one first neural network layer, and the input of the updated machine learning model includes the text and / or speech. Obtain prediction information corresponding to the first data, wherein the prediction information is generated by the updated machine learning model; Training is performed based on the expected information and the predicted information corresponding to the first data, and the goal of the training includes improving the similarity between the predicted information and the expected information.
11. The method according to claim 10, characterized in that, The first feature information is a feature map, and the feature map contains at least one location point. The method further includes: obtaining the location encoding information of each location point among the at least one location point. The step of updating the weight parameters of at least one first neural network layer in the machine learning model according to the first feature information includes: updating the weight parameters of the at least one first neural network layer in the machine learning model according to the first feature information and the position encoding information of each position point.
12. The method according to claim 11, characterized in that, The step of updating the weight parameters of at least one first neural network layer in the machine learning model based on the first feature information and the position encoding information of each location point includes: The first feature information is converted into second feature information by the first neural network module, and the second feature information is feature information in the form of a token. The weight parameters of at least one first neural network layer in the machine learning model are updated based on the second feature information and the location encoding information of each location point.
13. The method according to claim 12, characterized in that, The training based on the expected information and the predicted information corresponding to the first data includes: The first parameter is updated based on the expected information corresponding to the first data and the predicted information. The first parameter includes the weight parameters of the first neural network module and / or the position encoding information of each position point.
14. A data processing apparatus, characterized in that, The device includes: An acquisition module is used to acquire first data, the first data including an image, and the first data also including text and / or voice; The feature extraction module is used to extract features from the image to obtain the first feature information of the image; An update module is configured to update the weight parameters of at least one first neural network layer in a machine learning model based on the first feature information to obtain an updated machine learning model. The updated machine learning model includes the updated weight parameters of the at least one first neural network layer, and the input of the updated machine learning model includes the text and / or speech. The acquisition module is further configured to acquire prediction information corresponding to the first data, the prediction information being generated by the updated machine learning model.
15. A training device for a model, characterized in that, The device includes: An acquisition module is used to acquire first data, the first data including an image, and the first data also including text and / or voice; The feature extraction module is used to extract features from the image to obtain the first feature information of the image; An update module is configured to update the weight parameters of at least one first neural network layer in a machine learning model based on the first feature information to obtain an updated machine learning model. The updated machine learning model includes the updated weight parameters of the at least one first neural network layer, and the input of the updated machine learning model includes the text and / or speech. The acquisition module is further configured to acquire prediction information corresponding to the first data, the prediction information being generated by the updated machine learning model; A training module is used to train based on expected information and predicted information corresponding to the first data, wherein the training objective includes improving the similarity between the predicted information and the expected information.
16. An execution device, characterized in that, It includes a processor and a memory, wherein the processor is coupled to the memory. The memory is used to store programs; The processor is configured to execute a program in the memory, causing the execution device to perform the method as described in any one of claims 1 to 9.
17. A training device, characterized in that, It includes a processor and a memory, wherein the processor is coupled to the memory. The memory is used to store programs; The processor is configured to execute a program in the memory, causing the training device to perform the method as described in any one of claims 10 to 13.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 13.
19. A computer program product, characterized in that, The computer program product includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 13.