Model training method and related device
By adjusting the parameters of the feature extraction model to accurately express it on multiple information dimensions, the problem of insufficient expression ability of the feature extraction model in multi-dimensionality is solved, and the generalization and efficiency of feature extraction are improved.
Patent Information
- Application Number
- CN202410046659.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the feature extraction model has insufficient information expression ability in multiple information dimensions, resulting in poor generalization of features, difficulty in producing high-quality effects in multiple application fields, and low feature extraction efficiency.
By obtaining the feature processing model corresponding to the target feature extraction model and multiple information dimensions, the model parameters are adjusted using sample information sets, so that the feature extraction model can be accurately expressed on multiple information dimensions, and the generality and accuracy of information features are improved.
The accurate expression of the feature extraction model in multiple information dimensions is realized, which reduces the model training resources and time requirements, and improves the universality and efficiency of feature extraction.
Smart Images

Figure CN120298849A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning, and particularly to a model training method and related devices. Background Art
[0002] Feature extraction for information is one of the common functions of a model. Through the extracted features, various subsequent applications can be performed. For example, through the extracted image features, subsequent processing such as image content recognition, object contour segmentation, and image semantic recognition can be performed.
[0003] Among them, when performing different subsequent processes, it may depend on the information expression of features in different information dimensions. For example, when performing image semantic recognition, the information expressed by the image features in the semantic dimension is relied on. When performing object contour segmentation on the objects included in the image, the information expressed by the image features in the object shape dimension is relied on.
[0004] In the related art, when extracting features from information, the information expression ability of features in multiple information dimensions is not emphasized, resulting in poor versatility of the extracted features. It is difficult to produce high-quality feature application effects in multiple application fields, so feature extraction can only be performed separately for each application field, and the efficiency of feature extraction is low. Summary of the Invention
[0005] To solve the above technical problems, this application provides a model training method, which can strengthen the information expression of the information features extracted by the feature extraction model in each information dimension, thereby improving the versatility and accuracy of the information features.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, the embodiments of this application disclose a model training method, and the method includes:
[0008] Obtain a target feature extraction model and feature processing models respectively corresponding to multiple information dimensions. The target feature extraction model is used to extract information features corresponding to information, and the target feature processing model corresponding to the target information dimension is used to determine an information processing result according to the information expressed by the information features in the target information dimension. The information features are used to express the information expressed by the corresponding information in each information dimension;
[0009] Obtain the sample information sets corresponding to the multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has corresponding sample information processing results respectively. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information is any one of the sample information in the target sample information set;
[0010] Take the multiple information dimensions as the target information dimension respectively, and adjust the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model. The difference between the information processing result determined by the target feature processing model according to the first information feature and the target sample information processing result is less than a first preset threshold, where the first information feature is the information feature extracted by the feature extraction model according to the target sample information.
[0011] In a second aspect, an embodiment of the present application discloses a model training device, which includes a first acquisition unit, a second acquisition unit, and an adjustment unit:
[0012] The first acquisition unit is configured to acquire a target feature extraction model and feature processing models corresponding to multiple information dimensions respectively. The target feature extraction model is used to extract information features corresponding to information, and the target feature processing model corresponding to the target information dimension is used to determine an information processing result according to the information expressed by the information feature in the target information dimension. The information feature is used to express the information expressed by the corresponding information in each information dimension;
[0013] The second acquisition unit is configured to acquire the sample information sets corresponding to the multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has corresponding sample information processing results respectively. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information is any one of the sample information in the target sample information set;
[0014] The adjustment unit is configured to take the multiple information dimensions as the target information dimension respectively, and adjust the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model. The difference between the information processing result determined by the target feature processing model according to the first information feature and the target sample information processing result is less than a first preset threshold, where the first information feature is the information feature extracted by the feature extraction model according to the target sample information.
[0015] In a possible implementation, the target feature processing model is obtained through the following steps:
[0016] Obtain the initial feature processing model corresponding to the target information dimension;
[0017] Extract the second information feature corresponding to the target sample information through the initial feature extraction model, where the initial feature extraction model is used to extract the information feature corresponding to the information;
[0018] Determine the first information processing result corresponding to the second information feature through the initial feature processing model;
[0019] Adjust the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result to obtain the target feature processing model.
[0020] In a possible implementation, the step of adjusting the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result to obtain the target feature processing model includes:
[0021] Adjust the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result, so that the difference between the first information processing result and the target sample information processing result is less than a second preset threshold to obtain the target feature processing model, where the second preset threshold is greater than the first preset threshold.
[0022] In a possible implementation, the initial feature extraction model is the target feature extraction model.
[0023] In a possible implementation, the target feature extraction model is generated by adding adjustable parameters to the initial feature extraction model. The adjustable parameters are used to participate in extracting the information feature corresponding to the information, and the number of parameters of the adjustable parameters is less than the number of parameters included in the initial feature extraction model. The adjustment unit is specifically used for:
[0024] Respectively use the multiple information dimensions as the target information dimension, and adjust the adjustable parameters in the target feature extraction model through the target feature processing model and the target sample information set to obtain the feature extraction model.
[0025] In a possible implementation, the adjustment of the model parameters corresponding to the initial feature processing models respectively corresponding to the multiple information dimensions is executed in parallel.
[0026] In a possible implementation manner, the initial feature extraction model is trained as follows:
[0027] Obtain a feature extraction model to be trained and multiple pieces of sample information to be trained. The multiple pieces of sample information to be trained have respectively corresponding sample representation information, and the multiple pieces of sample representation information have respectively corresponding information features. The target sample representation information corresponding to the target sample information to be trained is used to represent the target sample information to be trained. The target sample representation information corresponds to a third information feature, and the feature extraction model to be trained is used to extract the information feature corresponding to the information;
[0028] Extract the to-be-determined information feature corresponding to the target sample information to be trained through the feature extraction model to be trained;
[0029] Adjust the feature extraction model to be trained according to the similarity between the to-be-determined information feature and the information features respectively corresponding to the multiple pieces of sample representation information, and obtain the initial feature extraction model. Among the information features respectively corresponding to the multiple pieces of sample representation information, the similarity between the to-be-determined information feature determined by the initial feature extraction model and the third information feature is the largest.
[0030] In a possible implementation manner, the sample information to be trained is image information, the sample representation information is text information, and the text information is used to express the information expressed by the corresponding image information.
[0031] In a possible implementation manner, the adjustment unit is specifically configured to:
[0032] Take the multiple information dimensions as the target information dimension respectively, and through the target feature extraction model, extract the first information feature according to the target sample information;
[0033] Determine the second information processing result corresponding to the first information feature through the target feature processing model;
[0034] Adjust the model parameters corresponding to the target feature extraction model according to the difference between the second information processing result and the target sample information processing result, and obtain the feature extraction model.
[0035] In a possible implementation manner, the sample information is image information, and the information processing results respectively corresponding to the multiple information dimensions include any combination of an object contour recognition result, an image semantic recognition result, an object category recognition result, and an image character recognition result.
[0036] In a possible implementation manner, the device further includes a third acquisition unit and a determination unit:
[0037] The third acquisition unit is configured to acquire a verification information set and a feature processing model respectively corresponding to multiple verification information dimensions. The target verification information set corresponding to the target verification information dimension includes multiple verification information, and the multiple verification information has respectively corresponding sample information processing results. The second sample information processing result corresponding to the target verification information is determined based on the information expressed by the target verification information in the target verification information dimension. The multiple verification information dimensions do not include the multiple information dimensions. The feature processing model corresponding to the target verification information dimension is configured to determine an information processing result according to the information expressed by the information feature in the target verification information dimension. The target verification information dimension is any one of the multiple verification information dimensions.
[0038] The determination unit is configured to verify the feature extraction model through the verification information sets and the feature processing models respectively corresponding to the multiple verification information dimensions, and determine the generalizability parameter corresponding to the feature extraction model. The generalizability parameter is used to characterize the accuracy of the information feature extracted by the feature extraction model in the information expressed by the multiple verification information dimensions. The accuracy characterized by the generalizability parameter is inversely correlated with the difference between the third information processing result and the second sample information processing result. The third information processing result is the information processing result determined by the feature processing model corresponding to the target verification information dimension according to the fourth information feature, and the fourth information feature is the information feature extracted by the feature extraction model according to the target verification information.
[0039] In a third aspect, an embodiment of the present application discloses a computer device, which includes a processor and a memory:
[0040] The memory is configured to store a computer program and transmit the computer program to the processor;
[0041] The processor is configured to execute the model training method according to any one of the items in the first aspect according to the instructions in the computer program;
[0042] In a fourth aspect, an embodiment of the present application discloses a computer-readable storage medium, which is configured to store a computer program, and the computer program is used to execute the model training method according to any one of the items in the first aspect;
[0043] In a fifth aspect, an embodiment of the present application discloses a computer program product including a computer program, which, when running on a computer device, causes the computer device to execute the model training method according to any one of the items in the first aspect.
[0044] As can be seen from the above technical solution, the present application provides a model training method. When training a feature extraction model for extracting information features corresponding to information, the target feature extraction model to be trained and feature processing models corresponding to multiple information dimensions can be obtained first. The target feature processing model corresponding to the target information dimension is used to determine an information processing result according to the information expressed by the information feature in the target information dimension, and the information feature is used to express the information expressed by the corresponding information in each information dimension. Moreover, sample information sets corresponding to multiple information dimensions can be obtained. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has corresponding sample information processing results respectively. The target sample information processing result corresponding to the target sample information is the accurate information processing result determined based on the information expressed by the target sample information in the target information dimension. Therefore, the difference between the information processing result determined by the target feature processing model based on the information feature extracted from the target sample information by the target feature extraction model and the target sample information processing result can characterize the accuracy and effectiveness of the information expression of the information feature extracted by the target feature extraction model in the target information dimension. Thus, multiple information dimensions can be used as the target information dimension respectively. By using the target feature processing model and the target sample information set, the model parameters corresponding to the target feature extraction model can be adjusted so that it learns how to strengthen the information expression of the extracted information feature in the target information dimension. Furthermore, through the feature processing models and sample information sets corresponding to multiple information dimensions respectively, the target feature extraction model can learn how to strengthen the information expression of the extracted information feature in multiple information dimensions, so that the obtained feature extraction model can be applicable to the information processing of multiple information dimensions, improving the versatility of the information feature extracted by the feature extraction model, reducing the model volume in the application of multiple information dimensions, and at the same time, there is no need to train the feature extraction model separately for multiple dimensions, reducing the model training time required in the application of multiple information dimensions. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 It is a schematic diagram of a model training method in an actual application scenario provided by an embodiment of the present application;
[0047] Figure 2 It is a flowchart of a model training method provided by an embodiment of the present application;
[0048] Figure 3 Schematic diagram of a model training method provided by an embodiment of the present application;
[0049] Figure 4 Flowchart of a model training method in an actual application scenario provided by an embodiment of the present application;
[0050] Figure 5 Schematic diagram of a model training method in an actual application scenario provided by an embodiment of the present application;
[0051] Figure 6 Schematic diagram of a model training method in an actual application scenario provided by an embodiment of the present application;
[0052] Figure 7 Schematic diagram of a model training method in an actual application scenario provided by an embodiment of the present application;
[0053] Figure 8 Schematic diagram of a model training method in an actual application scenario provided by an embodiment of the present application;
[0054] Figure 9 Structural block diagram of a model training device provided by an embodiment of the present application;
[0055] Figure 10 Structural diagram of a terminal provided by an embodiment of the present application;
[0056] Figure 11 Structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0057] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0058] The information dimension refers to the dimension for expressing information. Some information can be expressed in multiple information dimensions. For example, image information can express semantic information corresponding to the image information, the types of objects included in the image information, the corresponding object contours included in the image information, and other information. In the related art, the models for processing information are trained for a single information dimension. For example, in order to accurately identify semantic information from image information, an identification model for identifying semantic information can be trained. This identification model can extract information features from the image information and determine the semantic information corresponding to the image information based on the information features.
[0059] Therefore, the information features extracted by the models in the related art can only effectively express information for a single information dimension. When this information feature is applied to determine the information processing result corresponding to other information dimensions, since the information feature cannot effectively express the information for this information dimension, it is likely to result in an inaccurate determined information processing result. Thus, in the related art, in order to obtain the information processing results of information in multiple information dimensions, it is necessary to train the model in the information feature extraction part for each information dimension to extract effective information features. This requires consuming a large amount of model training resources, and when facing the information processing requirements of multiple information dimensions, multiple models need to be used, and the occupied running space and running resources are also relatively large.
[0060] To solve the above technical problems, the present application provides a model training method, which can obtain the feature processing models and sample information sets respectively corresponding to multiple information dimensions. The sample information corresponding to the target information dimension has a corresponding sample information processing result, and this sample information processing result is the accurate information processing result determined by the information expression of the sample information in the target information dimension. The feature processing model corresponding to the target information dimension is used to determine the information processing result in the target information dimension according to the information feature. Thus, through the difference between the sample information processing result and the information processing result actually determined by the information feature extracted based on the target feature extraction model, the effectiveness and accuracy of the information expression of the information feature extracted by the target feature extraction model in each information dimension can be analyzed. Furthermore, through the feature processing models and sample information sets respectively corresponding to multiple information dimensions, the parameters of the target feature extraction model used to extract information features can be adjusted, so that the trained feature extraction model can extract information features that can effectively express information in multiple information dimensions, thereby enabling the feature extraction model to be widely used in the information processing of multiple information dimensions, reducing the model training resources required for multi-information dimension applications, and improving the model training efficiency.
[0061] It can be understood that this method can be applied to a computer device, which is a computer device capable of performing model training, such as a terminal device or a server. This method can be independently executed by a terminal device or a server, or can be applied to a network scenario where a terminal device and a server communicate, and is executed in cooperation with the terminal device and the server. Among them, the terminal device can be a device such as a mobile phone, a tablet computer, a notebook computer, or a desktop computer. The terminal device can also include a variety of virtual reality devices, such as an augmented reality (AR) device, such as an AR glasses, an AR screen, etc., and can also include a virtual reality (VR) device, such as a head-mounted VR glasses. The server can be understood as an application server or a web server. In actual deployment, the server can be an independent server, a cluster server, or a cloud server, etc.
[0062] This application also relates to artificial intelligence (AI) technology. Artificial intelligence technology uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0063] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0064] This application mainly involves machine learning (ML) technology therein. Machine learning technology is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.
[0065] This application can complete the model training of multiple models through machine learning technology, such as the training of feature extraction models and feature processing models.
[0066] To facilitate the understanding of the technical solution provided by this application, next, a model training method provided by this application will be introduced in combination with an actual application scenario.
[0067] See Figure 1 , Figure 1 which is a schematic diagram of a model training method in an actual application scenario provided by an embodiment of this application. In this embodiment, the computer device may be a server 101 with model training capabilities.
[0068] The information may be image information, and multiple information dimensions may include a semantic recognition dimension and an object contour recognition dimension. Among them, the information processing result corresponding to the semantic recognition dimension is the semantic recognition result for the image information, and the information processing result corresponding to the object contour recognition dimension is the object contour recognition result, that is, the recognition result of the object contour of the object included in the image information.
[0069] The server 101 can obtain a feature processing model and a sample information set corresponding to two information dimensions respectively. Among them, the sample information set corresponding to the semantic recognition dimension includes three sample image information, namely sample image information 1, sample image information 2, and sample image information 3. The sample information set corresponding to the object contour recognition dimension includes three sample image information, namely sample image information 4, sample image information 5, and sample image information 6. Each sample image information has a corresponding sample information processing result. The sample information processing result of the semantic recognition dimension is the sample semantic recognition result, and the sample information processing result corresponding to the object contour recognition dimension is the sample object contour recognition result. The sample information processing result is the accurate information processing result.
[0070] Among them, the semantic recognition result is determined by the information expression of the feature processing model 1 in the semantic recognition dimension through information features, and the object contour recognition result is determined by the information expression of the feature processing model 2 in the object contour recognition dimension through information features. The target feature extraction model is used to extract the information features corresponding to the information, such as Figure 1 As shown, when training through two sample information sets, the sample image information in the two sample information sets can be first input into the target feature extraction model to extract the corresponding pending information features. Among them, the pending information feature corresponding to the sample image information pair 1 is the pending information feature 1, and the pending information feature corresponding to the sample image information 4 is the pending information feature 4. Through the feature processing model 1 corresponding to the semantic recognition dimension, the pending semantic recognition result can be determined based on the pending information feature 1. Through the feature processing model 2 corresponding to the object contour recognition dimension, the pending object contour recognition result can be determined based on the pending information feature 4. Since the feature processing model 1 and the feature processing model 2 itself have accurate feature processing capabilities, therefore, whether the output information processing result is accurate largely depends on whether the information features can have accurate information expressions in the corresponding information dimensions.
[0071] Thus, through the difference between the pending semantic recognition result and the sample semantic recognition result 1 corresponding to the sample image information 1, it can be characterized whether the information expression of the pending information feature 1 extracted by the target feature extraction model in the semantic recognition dimension is accurate. Through the difference between the pending object contour recognition result and the sample object contour recognition result 4 corresponding to the sample image information 4, it can be characterized whether the information expression of the pending information feature 4 extracted by the target feature extraction model in the object contour recognition dimension is accurate. Based on this, through the sample information sets and feature processing models corresponding to these two information dimensions, the accuracy of the information expression of the information features extracted by the target feature extraction model in each information dimension can be analyzed. Thus, the model parameters of the target feature extraction model can be adjusted to obtain a feature extraction model, so that the feature extraction model satisfies that the information processing results corresponding to the information features extracted in each information dimension can have a high degree of accuracy, so that the information features extracted in each information dimension have accurate information expressions, thereby improving the versatility of the information features extracted by the feature extraction model. When facing information processing applications in multiple information dimensions, only this feature extraction model is needed to meet all feature extraction requirements, without the need to configure a separate feature extraction model for each information dimension, reducing the resource amount required for model training and improving the model training efficiency.
[0072] Next, the technical solution provided by this application will be introduced in detail with reference to the accompanying drawings.
[0073] See Figure 2 , Figure 2The flowchart of a model training method provided by an embodiment of the present application. In this embodiment, the computer device can be any of the above computer devices with model training functions, including:
[0074] S201: Obtain a target feature extraction model and feature processing models corresponding to multiple information dimensions.
[0075] Among them, the feature processing model is a pre-trained model with accurate feature processing capabilities. The feature processing capability refers to the ability to determine an information processing result based on information features. The information dimension is the dimension for expressing information, and the information processing results corresponding to different information dimensions are different, that is, the results obtained by performing information processing based on the information expressed in different information dimensions are different. The multiple information dimensions can be any multiple information dimensions in which the same information can be expressed. For example, in Figure 1 , the information is image information, and the multiple information dimensions can include a semantic recognition dimension and an object contour recognition dimension. The information processing result corresponding to the semantic recognition dimension is a semantic recognition result, and the information processing result corresponding to the object contour recognition dimension is an object contour recognition result.
[0076] The target feature extraction model is used to extract information features corresponding to the information. The target feature processing model corresponding to the target information dimension is used to determine the information processing result based on the information expressed by the information features in the target information dimension. The information features are used to express the information expressed by the corresponding information in each information dimension. The target information dimension can be any one of the multiple information dimensions. That is, if the information features extracted by the target feature extraction model can accurately express information in multiple information dimensions, the feature processing models corresponding to the multiple information dimensions can all obtain accurate information processing results based on this information feature. For example, if Figure 1 the undetermined information features determined by the target feature extraction model in the semantic recognition dimension can have accurate information expression, the feature processing model 1 can determine an accurate semantic recognition result based on this undetermined information feature.
[0077] It should be emphasized that the models involved in the present application are not limited to specific model structures and model types, as long as they have the model functions required by the present application.
[0078] S202: Obtain sample information sets corresponding to multiple information dimensions respectively.
[0079] In order to measure whether the information features extracted by the target feature extraction model can accurately express information in multiple information dimensions during model training, the computer device can obtain the sample information sets corresponding to multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has corresponding sample information processing results respectively. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information can be any one of the sample information in the target sample information set.
[0080] That is, the sample information processing result is an information processing result that can accurately correspond to the information expression of the sample information in the corresponding information dimension. If the information features extracted by the target feature extraction model based on the target sample information can accurately express information in the target information dimension, the information processing result determined based on this information feature by the target feature processing should be close to the target sample information processing result. Therefore, based on the sample information processing result, the accuracy of the information expression of the information features extracted by the target feature extraction model in each information dimension can be analyzed.
[0081] S203: Use multiple information dimensions as the target information dimension respectively, and adjust the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain the feature extraction model.
[0082] The computer device can use the target feature processing model to determine the information processing result corresponding to the information features extracted by the target feature extraction model in the target information dimension, and use the sample information processing results corresponding to the sample information in the target sample information set to measure whether the information features can accurately express information in the target information dimension based on the above method. Therefore, based on this measurement result, when accurate information expression cannot be performed, the model parameters of the target feature extraction model can be adjusted so that it learns how to extract information features that can accurately express information in the target information dimension.
[0083] Thus, based on the feature processing models corresponding to multiple information dimensions respectively and the sample information set, the parameter adjustment of multiple information dimensions can be jointly reflected in the parameter adjustment of the target feature extraction model. Furthermore, the target feature extraction model can learn how to extract information features that can accurately express information in multiple information dimensions, and obtain a trained feature extraction model. It can enable the difference between the information processing result determined by the target feature processing model based on the first information feature and the target sample information processing result to be less than a first preset threshold, where the first information feature is the information feature extracted by the feature extraction model based on the target sample information. The first preset threshold is used to measure whether the two information processing results are close enough, and further to measure whether the information processing result determined by the target feature processing model is accurate. If the difference between the information processing result and the target sample information processing result is less than the first preset threshold, it indicates that the information processing result determined by the target feature processing model based on the first information feature is relatively accurate, thus indicating that the information feature extracted by the feature extraction model based on the target sample information can accurately express information in the target information dimension.
[0084] As can be seen from the above technical solution, the present application provides a model training method, which can enable the feature extraction model to learn how to strengthen the information expression of the extracted information features in multiple information dimensions. Thus, the obtained feature extraction model can be applicable to information processing in multiple information dimensions, improving the versatility of the information features extracted by the feature extraction model. In the face of application scenarios with multiple information dimensions, only this one feature extraction model is needed to meet the feature extraction requirements, reducing the model volume in the case of multiple information dimensions application. At the same time, there is no need to train the feature extraction model separately for multiple dimensions, reducing the model training time required for multiple information dimensions application. Meanwhile, when a new information dimension appears, since the information features extracted in the present application are more comprehensive, on the one hand, good information processing results can also be determined in the new information dimension. On the other hand, if it is necessary to train a targeted feature extraction model for the new information dimension, training based on the feature extraction model determined in the present application can also accelerate the model fitting speed, laying a good training foundation for the new information dimension.
[0085] Next, each link of the model training in the present application will be introduced in detail.
[0086] First, as mentioned above, the feature processing models corresponding to multiple information dimensions obtained in step S201 have accurate feature processing capabilities, which is also the key for the present application to directly transmit the difference in information processing results back to the target feature extraction model. Next, it will be introduced how to enable the feature processing model to learn accurate feature processing capabilities.
[0087] Taking the target feature processing model as an example, in a possible implementation, the target feature processing model can be obtained in the following way:
[0088] The computer device can first obtain an initial feature processing model corresponding to the target information dimension. This initial feature processing model is a feature processing model that has not been trained yet. Then, the computer device can extract the second information feature corresponding to the target sample information through the initial feature extraction model. The initial feature extraction model is used to extract the information feature corresponding to the information. Here, the initial feature extraction model only needs to have the feature extraction ability, and the model relationship with the target feature extraction model is not limited, which will be introduced in detail later.
[0089] Then, the computer device can determine the first information processing result corresponding to the second information feature through the initial feature processing model. Since the target sample information processing result corresponding to the target sample information is the accurate information processing result corresponding to the target sample information in the target information dimension, and the first information processing result is the information processing result actually output by the initial feature processing model during training, the difference between the first information processing result and the first sample information processing result can reflect the accuracy of the initial feature processing model for feature processing based on the information feature, as well as the accuracy of the initial feature extraction model for feature extraction. In this embodiment, in order to enable the initial feature processing model to learn as much as possible how to accurately perform feature processing, the object of parameter adjustment can be concentrated on the initial feature processing model, that is, only according to the difference between the first information processing result and the target sample information processing result, adjust the model parameters corresponding to the initial feature processing model to obtain the target feature processing model, so as to increase the learning intensity of the initial feature processing model, ensure that the target feature processing model has a high feature processing accuracy, and then when training the target feature extraction model based on this target feature processing model, it can ensure high-precision adjustment of the model parameters of the target feature extraction model, thereby improving the accuracy of model training.
[0090] It should be emphasized that the above training process for the feature processing model uses the same sample information set as the training feature extraction model. The computer device can also use different sample information sets to train the feature processing model, which is not limited here.
[0091] It can be understood that although the model accuracy of the feature processing model can be improved by the above method, since the model parameters of the initial feature extraction model are not adjusted, if the training accuracy requirement for the feature processing model is too high, it may lead to the problem of overfitting of the feature processing model. Based on this, in a possible implementation, when adjusting the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result to obtain the target feature processing model, the computer device can adjust the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result, so that the difference between the first information processing result and the target sample information processing result is less than a second preset threshold to obtain the target feature processing model, and the second preset threshold is greater than the first preset threshold.
[0092] Thus, on the one hand, since the second preset threshold is smaller than the first preset threshold, the accuracy of the information processing result that needs to be improved by adjusting the initial feature extraction model can be hidden, avoiding the initial feature processing model from overlearning this part of the model knowledge that it should not learn, and reducing the probability of overfitting; on the other hand, since the first preset threshold is higher than the second preset threshold, during the process of training the feature extraction model based on the first preset threshold, the feature extraction model can successfully learn this part of the model knowledge, avoiding the situation where the feature extraction model cannot learn through the feedback of the difference due to the high feature processing accuracy of the feature processing model itself resulting in the output information processing result meeting the accuracy requirement.
[0093] Combined with the above content, it can be seen that in the training stage of the feature processing model of this application, the adjustment of the model parameters of the feature extraction model is blocked, and only the model parameters of the feature processing model are adjusted. Firstly, it ensures that the feature processing model has accurate feature processing capabilities, which can be used as the training basis of the feature extraction model on the one hand, and on the other hand, the feature processing model itself can also be independently applied in various application scenarios; then, in the training stage of the feature extraction model, the adjustment of the model parameters of the feature processing model is blocked, and only the model parameters of the feature extraction model are adjusted, so that the feature extraction knowledge of multiple information dimensions can be accurately transmitted to the feature extraction model, realizing the effective training of the feature extraction model, and ensuring that the information features extracted by the feature extraction model can accurately express information in multiple information dimensions.
[0094] As mentioned above, the model relationship between the initial feature extraction model and the target feature extraction model can include various types, which will be introduced in detail below.
[0095] First, in a possible implementation, the initial feature extraction model can be the target feature extraction model. The advantage of this relationship is that since the same feature extraction model participates in the training stage of the feature processing model, the way the feature processing model processes information features is more in line with the characteristics of the information features extracted by the target feature extraction model. Thus, it can be avoided that the target feature extraction model adjusts its model parameters to fit the feature processing model, reducing the time required for the target feature extraction model to fit and further improving the model training efficiency.
[0096] Secondly, it can be understood that since the feature extraction model of the present application is required to extract information features with accurate information expressions in multiple information dimensions, to ensure the accuracy of information feature extraction, a large number of model parameters for feature extraction are usually included in the feature extraction model. If all the parameters of the feature extraction model are adjusted, a large amount of training resources and time will be consumed.
[0097] Based on this, in another possible implementation, to further improve the model training efficiency while ensuring the model training quality, the target feature extraction model can be generated by adding adjustable parameters to the initial feature extraction model. The adjustable parameters are used to participate in extracting the information features corresponding to the information, and the number of parameters of the adjustable parameters is less than the number of parameters included in the initial feature extraction model.
[0098] When performing step S203, the computer device can execute step S2031 (not shown in the figure). Step S2031 is a possible implementation of step S203 and includes:
[0099] S2031: Use multiple information dimensions as target information dimensions respectively, and through the target feature processing model and the target sample information set, adjust the adjustable parameters in the target feature extraction model to obtain a feature extraction model.
[0100] That is, when performing parameter adjustment, the computer device can only adjust this part of the newly added adjustable parameters, making the learning of the information expressions in multiple information dimensions focus on the adjustable parameters. Since the number of parameters of the adjustable parameters is less than the number of parameters of the initial feature extraction model, the range of parameter adjustment is much smaller than directly adjusting all the model parameters of the initial feature extraction model, accelerating the fitting of the model parameters and improving the efficiency of model training. In addition, since the adjustable parameters are used to participate in extracting information features, it can be ensured that adjusting the adjustable parameters enables the target feature extraction model to effectively learn how to accurately extract information features, thus ensuring the quality of model training.
[0101] In addition to improving the model efficiency through the dimensions of model parameters, the computer device can also improve the model training efficiency from the overall model training process. For example, as can be seen from the above, during the process of training the feature processing model, the model parameters of the initial feature extraction model are locked and unchanged. Therefore, when each information processing model corresponding to an information dimension is trained, the initial feature extraction model used can be the same model. Based on this, in a possible implementation manner, the adjustment of the model parameters of the initial feature processing models corresponding to multiple information dimensions is performed in parallel. That is, the computer device can make the initial feature extraction model output the information features corresponding to multiple information dimensions simultaneously, or the computer device can prepare multiple initial feature extraction models with the same parameters, and each initial feature extraction model is used for the training of the information processing model of one information dimension. Thus, the training duration of the feature processing models of N information dimensions can be reduced from N times to at least 1 time, significantly reducing the time required for model training.
[0102] As mentioned above, the initial feature extraction model is a model with certain feature extraction capabilities, and there can also be various training methods for the initial feature extraction model. Next, a unsupervised training method will be mainly introduced.
[0103] In a possible implementation manner, the initial feature extraction model can be trained through the following method:
[0104] The computer device can first obtain the feature extraction model to be trained and multiple pieces of information of samples to be trained. The feature extraction model to be trained is the initial feature extraction model before training. The multiple pieces of information of samples to be trained have respectively corresponding sample representation information, and the multiple sample representation information have respectively corresponding information features. The target sample representation information corresponding to the target information of the sample to be trained is used to represent the target information of the sample to be trained, and the target sample representation information corresponds to the third information feature. The feature extraction model to be trained is used to extract the information features corresponding to the information. Among them, the target information of the sample to be trained can be any one of the multiple pieces of information of samples to be trained.
[0105] Since the target sample representation information can represent the information of the target sample to be trained, the expression of the target sample representation information and the information of the target sample to be trained in the information dimension has a high similarity. Therefore, the information features corresponding to the two should also have a high similarity. Based on this, the computer device can extract the to-be-determined information features corresponding to the information of the target sample to be trained through the to-be-trained feature extraction model. If the to-be-determined information features can accurately represent the information expression of the target sample to be trained in each information dimension, the to-be-determined information features should have a high similarity with the third information features and a low similarity with the information features of the sample representation information corresponding to other samples to be trained. Based on this, the computer device can adjust the to-be-trained feature extraction model according to the similarity between the to-be-determined information features and the information features corresponding to multiple sample representation information, and obtain an initial feature extraction model, so that among the information features corresponding to multiple sample representation information, the similarity between the to-be-determined information features determined by the initial feature extraction model and the third information features is the largest. Thus, the initial feature extraction model can learn how to make the determined information features represent the information expressed in the information dimension of the corresponding information.
[0106] Among them, based on the different information types of the samples to be trained, the information types of the corresponding sample representation information can also be different. For example, in a possible implementation, the sample information to be trained can be image information, and the sample representation information is text information, which is used to express the information expressed by the corresponding image information. For example, the image information can be a photo taken of a puppy playing with a ball, and the text information can be "The puppy is playing with a ball", so that these two pieces of information can have a relatively similar information expression, and thus have a high similarity in information features.
[0107] Next, the model training process in step S203 will be introduced in more detail.
[0108] In a possible implementation, when executing step S203, the computer device can execute steps S2032 - S2034 (not shown in the figure), and steps S2032 - S2034 are a possible implementation of step S203, including:
[0109] S2032: Take multiple information dimensions as target information dimensions respectively, and through the target feature extraction model, extract the first information features according to the target sample information.
[0110] The first information features can represent the information expression of the target sample information analyzed by the target feature extraction model in multiple information dimensions.
[0111] S2033: Determine the second information processing result corresponding to the first information features through the target feature processing model.
[0112] The target feature processing model can determine the second information processing result based on the information expression of the first information feature in the target information dimension.
[0113] S2034: Adjust the model parameters corresponding to the target feature extraction model according to the difference between the second information processing result and the target sample information processing result to obtain a feature extraction model.
[0114] Since the target feature processing model has the ability to accurately determine the information processing result based on the information expressed in the target information dimension, the accuracy of the information processing result determined by the target feature processing model mainly depends on the accuracy of the information expression of the information features extracted by the target feature processing model in the target information dimension. That is, the more accurate the information expression of the first information feature in the target information dimension, the closer the second information processing result determined by the target feature processing model should be to the target sample information processing result. Thus, the difference between the second information processing result and the target sample information processing result can characterize the accuracy of the feature extraction by the target feature extraction model. The computer device can adjust the model parameters of the target feature extraction model according to this difference, so that the second information processing result determined based on this model gradually approaches the target sample information processing result. In this process, the target feature extraction model can learn how to strengthen the information expression of the extracted information features in the target information dimension. Furthermore, through the overall training process involved in step S203, the target feature extraction model can learn how to strengthen the information expression of the extracted information features in multiple information dimensions to obtain the required feature extraction model.
[0115] As mentioned above, the information involved in this application can include multiple information types, and the information dimensions corresponding to different information types may also be different. Next, the information dimensions involved in image information will be mainly introduced in detail.
[0116] In a possible implementation manner, the sample information can be image information, and the information processing results corresponding to multiple information dimensions can include any combination of an object contour recognition result, an image semantic recognition result, an object category recognition result, an image character recognition result, etc.
[0117] Among them, the object contour recognition result is the result of recognizing the object contour of the object included in the image information, the image semantic recognition result is the result of recognizing the semantic information expressed by the image information, the object category recognition result is the result of recognizing the object category of the object included in the image information, and the image character recognition result is the result of recognizing the characters included in the image information. Of course, other multiple information processing results can also be included for image information, which is not limited here.
[0118] Such asFigure 3 As shown, the image information can record the image information of a dog playing with a frisbee. Among them, through the object type recognition result corresponding to the image information, it can be determined that the image information includes two objects, namely, "dog" and "frisbee". The object contour recognition result can include the recognition result of the dog's contour and the recognition result of the frisbee's contour. The image semantic recognition result can be the text information "A dog is playing with a frisbee", and this text information can express the semantics corresponding to the image information.
[0119] From the above content, it can be seen that the purpose of this application is to enable the feature extraction model to strengthen the information expression of the extracted information features in multiple information dimensions and improve the versatility of the information features. Therefore, in a possible implementation manner, in order to test the effect of model training, the computer device can introduce new information dimensions to measure the versatility of the information features extracted by the feature extraction model.
[0120] The computer device can obtain the verification information sets and feature processing models corresponding to multiple verification information dimensions respectively. Similar to the above-mentioned multiple information dimensions, the target verification information set corresponding to the target verification information dimension includes multiple verification information, and the multiple verification information has corresponding sample information processing results respectively. The sample information processing result is the accurate information processing result corresponding to the verification information. The target verification information dimension can be any one of the multiple verification information dimensions. The second sample information processing result corresponding to the target verification information is determined based on the information expressed by the target verification information in the target verification information dimension. That is, if the information features extracted by the feature extraction model based on the target verification information can have an accurate information expression in the target verification information dimension, the information processing result determined based on the information features extracted by the feature extraction model should be relatively close to this second sample information processing result. The target verification information can be any one of the verification information included in the target verification information set.
[0121] It should be emphasized that the multiple verification information dimensions do not include the multiple information dimensions because the multiple information dimensions have enabled the model to be fully learned in the above training process. The purpose of this implementation manner is to test the versatility of the information feature model in more information dimensions. The feature processing model corresponding to the target verification information dimension can be used to determine the information processing result according to the information expressed by the information features in the target verification information dimension.
[0122] A computer device can verify a feature extraction model through a set of verification information corresponding to multiple verification information dimensions and a feature processing model, and determine a general-purpose parameter corresponding to the feature extraction model. This general-purpose parameter is used to characterize the accuracy of the information features extracted by the feature extraction model in expressing information in multiple verification information dimensions. The higher the accuracy, the wider the information dimension application of the information features extracted by the feature extraction model, and the higher the generality. Specifically, taking the target verification information dimension as an example, the computer device can input the target verification information into the feature extraction model to extract the corresponding information features, and then input these information features into the feature processing model corresponding to the target verification information dimension to obtain the corresponding information processing result. The higher the similarity between this information processing result and the second sample information processing result, the more accurately the information features extracted by the feature extraction model can express information in the target verification information dimension. Therefore, it can be shown that the information features extracted by the feature extraction model can be effectively applied even in information dimensions not involved in the training process, and have strong generality. That is, the accuracy characterized by this general-purpose parameter is inversely correlated with the difference between the third information processing result and the second sample information processing result. The third information processing result is the information processing result determined by the feature processing model corresponding to the target verification information dimension according to the fourth information feature, and the fourth information feature is the information feature extracted by the feature extraction model according to the target verification information. The smaller this difference, the more accurately the information features extracted by the feature extraction model express information in the target verification information dimension. The more accurately the information is expressed in multiple verification information dimensions, the stronger the generality of the extracted information features.
[0123] To facilitate the understanding of the technical solution provided by this application, next, a model training method provided by this application will be introduced in combination with an actual application scenario.
[0124] See Figure 4 , Figure 4 is a flowchart of a model training method in an actual application scenario provided by an embodiment of this application. In this actual application scenario, the information can be image information, and multiple information dimensions include a semantic recognition dimension, an object contour recognition dimension, and an object type recognition dimension. The computer device can be any computer device with model training capabilities, such as a computer or a server used for model training. The method includes:
[0125] S401: Train an initial feature extraction model.
[0126] The computer device can train the initial feature extraction model based on the similarity of information features. For example, Figure 5As shown, the feature extraction model to be trained can be a mainstream image + language multi-modal model (Contrastive Language-Image Pre-training, abbreviated as CLIP). This model can include two Add&Norm layers, a Feed Forward layer, and a Multi-Head Self Attention layer. After the image information is input into the feature extraction model to be trained, it will first perform feature extraction through the Multi-Head Self Attention layer. The Multi-Head Self Attention layer is a variant based on the Attention layer. The Multi-Head Self Attention layer introduces the "multi-head" mechanism on the basis of the Attention layer, splits the input into multiple sub-spaces, each sub-space separately executes the functions of the Attention layer, and finally splices the outputs of multiple sub-spaces together and performs a linear transformation to obtain the final output.
[0127] The input of the Multi-Head Self Attention layer will enter the first Add&Norm layer. The Add&Norm layer consists of two parts: Add and Norm. Here, Add refers to X + the Multi-Head Self Attention layer, which is a residual connection. Norm is the Layer Normalization layer, which normalizes the input. After passing through the first Add&Norm layer, it will enter the Feed Forward layer. The full name of the Feed Forward layer is Position-wise Feed-Forward Networks, and its essence is a two-layer fully connected layer. The activation function of the first layer is the Linear rectification function (abbreviated as Relu), and the second layer does not use an activation function. Then, after passing through the second Add&Norm layer, the image information features of the output can be obtained.
[0128] By the similarity between the text information features extracted based on text information and the image information features extracted through the above process, the parameters of the feature extraction model to be trained can be adjusted to obtain an initial feature extraction model. The similarity between the image information features corresponding to the image information extracted by this initial feature extraction model and the text information features of the corresponding text information is the highest, and the similarity with the text information features of other text information is lower.
[0129] S402: Lock the model parameters corresponding to the initial feature extraction model, and train to obtain feature processing models corresponding to multiple information dimensions through the initial feature extraction model and the sample information sets corresponding to multiple information dimensions.
[0130] Such as Figure 6As shown, in this application, multiple information dimensions may include a semantic recognition dimension, an object category recognition dimension, and an object contour recognition dimension. Among them, the feature processing models selected for each information dimension may be as follows:
[0131] For the information dimension of object category recognition, this application may use a detection model (DEtectionTransformer, abbreviated as Detr) as the feature processing model. Detr generates a fixed number of learnable parameters as the input to the image decoder. These learnable parameters influence each other through self-attention and interact with the flattened image features through cross-attention layers. Subsequently, a Multilayer Perceptron (MLP) and a linear head are respectively used for bounding box and classification label prediction. Finally, a bipartite graph matching mechanism is used to match the prediction results with the ground truth boxes, so as to obtain the categories of each object in the image information and the region boxes where the objects are located.
[0132] For the information dimension of object contour recognition, this application may use an image segmentation model (Mask2former) as the feature processing model. Mask2former also generates a fixed number of learnable parameters. The segmentation mask comes from the dot product between the hidden state of the i-th learnable parameter in the final layer of the decoder and each pixel feature map:
[0133]
[0134] where is a 1x1 convolutional layer followed by group normalization (GN); is a 1x1 convolution, followed by group normalization and bilinear upsampling; is a 3x3 convolution, followed by GN, ReLU, and a 1x1 convolution. and respectively represent the per-pixel feature maps generated by the backbone network and the segmentation head. Through the above process, the object contours corresponding to each object in the image information can be finally output.
[0135] For the information dimension of semantic recognition, this application may use a Long Short Term Memory (LSTM) network, which generates the semantic recognition result of the image by generating a word at each time step, conditioned on the context vector, the previous hidden state, and the previously generated words.
[0136] The computer device may first obtain a sample information set D(x, y), where x is the sample image information and y is the corresponding sample information processing result. Then, initialize the feature processing models T corresponding to multiple information dimensions respectively n, n ∈ {1, 2, 3}, where 1, 2, and 3 represent the semantic information dimension, the object type recognition dimension, and the object contour recognition dimension respectively. After keeping the model parameters of the initial feature extraction model unchanged, the feature processing models corresponding to multiple information dimensions can be trained in parallel. Among them, after inputting the sample image information, the corresponding image feature information f = M(x) can be obtained through the initial feature extraction model, where M is the initial feature extraction model. By minimizing the loss function L n (y, T n (f)) as the adjustment objective of the model parameters, the feature processing models corresponding to each information dimension can be obtained, where T n (f) is the information processing result output by the initial feature processing model.
[0137] S403: Add adjustable parameters to the initial feature extraction model to obtain the target feature extraction model.
[0138] As Figure 7 shown, this application can use the low-rank adaptation (LoRA) as the adjustable parameter, which can act on the multi-head attention mechanism layer and jointly affect the output of the multi-head attention mechanism layer with the weight parameters (such as W q / v ) of the pre-trained initial feature extraction model.
[0139] S404: Block the model parameters corresponding to the feature processing model, and train the target feature extraction model through the sample information sets and feature processing models corresponding to multiple information dimensions to obtain the feature extraction model.
[0140] During the training process, keep the model parameters of the multiple trained feature processing models unchanged, and select the sample image information in the sample information sets corresponding to multiple information dimensions based on the sampling ratios α n corresponding to multiple information dimensions respectively, so that the number of sample image information of each finally selected information dimension meets the sampling ratio. Train the model through these sample image information, and finally adjust the LoRA parameter ΔW based on the minimization of the loss function L′ n (y, T n (f′)) to obtain the adjusted LoRA parameter ΔW * , and then obtain the feature extraction model. Among them, f′ is the information feature output by the target feature extraction model, and T n (f′) is the information processing result output by the feature processing model corresponding to the information dimension n.
[0141] S405: Verify the feature extraction model through the feature processing models and verification information sets corresponding to multiple verification information dimensions, and determine the corresponding generalization parameters.
[0142] See Figure 8 This application can obtain the feature processing models T corresponding to O verification information dimensions respectively O and the sample information sets E corresponding to these verification information dimensions respectively O (x, y), O ∈ {1, …, O}, x is the sample image information, y is the corresponding human sample information processing result. The image information feature f = M(x; ΔW) is extracted through the trained feature extraction model M, and the information processing result T O (f) is obtained based on this information feature. Finally, the generality parameters corresponding to each verification information dimension can be obtained through the difference between this information processing result and the sample information processing result, and the generality parameter Metric(y, T O (f)) corresponding to this feature extraction model is comprehensively obtained, where Metric is used to measure the distance between two information processing results. The smaller the distance, the higher the accuracy indicated by the generality parameter.
[0143] For example, this application can verify the improvement effect of the technical solution on 6 different verification information dimensions and sample image information sets:
[0144] (1) This application can be tested on different Optical Character Recognition (OCR) data sets, and the average accuracy is reported. The results in Table 1 show that after applying our technical solution, the performance of optical character recognition has increased by at least 2.5 points, indicating that the feature extraction model of this application effectively improves the fine-grained information expression of information features in the optical character recognition dimension after training and can capture the complex details and semantic information of images.
[0145]
[0146] Table 1 Improvement effect of the feature extraction model on OCR after applying the technical solution of this application
[0147] (2) In Table 2, the effectiveness and robustness of this application in 9 zero-shot image classification benchmarks are further demonstrated. We conducted experiments on the EVA-CLIP-E model and observed improvements in all 9 data sets. On data sets composed of adversarial and unmodified examples, such as the ImageNet-A data set (from 82.1% to 82.4%) and the EuroSAT data set (from 65.8% to 67.1%), the enhancement effect is obvious, indicating that this solution can enhance the robustness of the model to real-world perturbations.
[0148]
[0149] Table 2 Improvement effect of the visual foundation model in zero-shot image classification after applying the technical solution of this patent
[0150] (3) Specified object recognition (GOI) involves classifying the specified object in an image using the information features extracted by the feature extraction model. As shown in Table 3, both the EVA-ViT-G model and the EVA-ViT-E model have improvements ranging from 0.3 to 0.6 on the M3IT dataset.
[0151]
[0152] Table 3 Improvement effect of the visual foundation model in specified object recognition after applying the technical solution of this patent
[0153] (4) Table 4 lists the zero-shot image and text retrieval results on Flickr30K and COCO. After implementing our technical solution, the EVA-CLIP-E model has improvements in both text and image retrieval, with the impact on the image retrieval task being more significant. For example, evaluated according to the Recall@5 metric of COCO, the image retrieval performance of EVA-CLIP-E has increased by 1.1%. This is due to the fact that this solution enables the feature extraction model to better understand and extract the relevant features paired with the corresponding text in the image.
[0154]
[0155] Table 4 Improvement effect of the visual foundation model in zero-shot image-text retrieval after applying the technical solution of this patent
[0156] The present invention also further evaluated examples other than EVA-CLIP (such as BLIP-2). As shown in Table 4, after fine-grained adjustment of the visual encoder of BLIP-2, an optimization phenomenon similar to that of EVA-CLIP can be observed.
[0157] (5) This technical solution also evaluated the zero-shot visual question answering performance of BLIP-2 ViT-G OPT 2.7B and BLIP-2 ViT-G OPT 6.7B on benchmarks such as VQAv2, GQA, and OK-VQA. As shown in Table 5, the models that perform ViSFT (i.e., the training method for training the feature extraction model in this application) on the visual encoder maintain or improve their performance in all three benchmark tests. The improvement on OK-VQA is more obvious, indicating that this solution can bring advantages to out-of-domain datasets.
[0158]
[0159]
[0160] Table 5 Improvement effect of the feature extraction model on zero-shot visual question answering after applying the technical solution of this patent
[0161] (6) This patent also evaluated BLIP-2 ViT-G OPT 2.7B The image annotation performance on the NoCaps dataset after using the technical solution of this patent. This dataset was not used during the training phase. The research results show that the technical solution of this patent can improve the annotation performance on unseen datasets. The results are shown in Table 6
[0162]
[0163] Table 6 Improvement effect of the feature extraction model on image annotation on the NoCaps dataset after applying the technical solution of this patent
[0164] Thus, in practical applications, the training method of this application can enable the information features extracted by the trained feature extraction model to have accurate information expressions in multiple information dimensions, so that the information features can be applied to various information processing scenarios without the need to train a targeted feature extraction model for each information processing scenario. While ensuring the effectiveness of the information features, it greatly reduces the time required for model training and improves the versatility of the feature extraction model
[0165] Based on the model training method provided in the above embodiments, the embodiments of this application also provide a model training device. See Figure 9 , Figure 9 which is the structural block diagram of a model training device provided by the embodiments of this application. The device 900 includes a first acquisition unit 901, a second acquisition unit 902, and an adjustment unit 903
[0166] The first acquisition unit 901 is used to acquire a target feature extraction model and feature processing models corresponding to multiple information dimensions respectively. The target feature extraction model is used to extract information features corresponding to information. The target feature processing model corresponding to the target information dimension is used to determine an information processing result according to the information expressed by the information feature in the target information dimension. The information feature is used to express the information expressed by the corresponding information in each information dimension
[0167] The second acquisition unit 902 is used to acquire sample information sets corresponding to the multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information. The multiple sample information has respectively corresponding sample information processing results. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information is any sample information in the target sample information set
[0168] The adjustment unit 903 is configured to use each of the multiple information dimensions as the target information dimension, and adjust the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set, so as to obtain a feature extraction model. The difference between the information processing result determined by the target feature processing model based on the first information feature and the target sample information processing result is less than a first preset threshold, and the first information feature is the information feature extracted by the feature extraction model from the target sample information.
[0169] In a possible implementation manner, the target feature processing model is obtained through the following steps:
[0170] Obtain an initial feature processing model corresponding to the target information dimension;
[0171] Extract a second information feature corresponding to the target sample information through an initial feature extraction model, where the initial feature extraction model is used to extract the information feature corresponding to the information;
[0172] Determine a first information processing result corresponding to the second information feature through the initial feature processing model;
[0173] Adjust the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result, so as to obtain the target feature processing model.
[0174] In a possible implementation manner, the adjusting the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result to obtain the target feature processing model includes:
[0175] Adjust the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result, so that the difference between the first information processing result and the target sample information processing result is less than a second preset threshold, and the target feature processing model is obtained, where the second preset threshold is greater than the first preset threshold.
[0176] In a possible implementation manner, the initial feature extraction model is the target feature extraction model.
[0177] In a possible implementation manner, the target feature extraction model is generated by adding adjustable parameters to the initial feature extraction model. The adjustable parameters are used to participate in extracting the information feature corresponding to the information, and the number of parameters of the adjustable parameters is less than the number of parameters included in the initial feature extraction model. The adjustment unit 903 is specifically configured to:
[0178] Taking the multiple information dimensions as the target information dimensions respectively, and adjusting the parameters to be adjusted in the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model.
[0179] In a possible implementation, the adjustment of the model parameters of the initial feature processing models corresponding to the multiple information dimensions is performed in parallel.
[0180] In a possible implementation, the initial feature extraction model is trained in the following manner:
[0181] Obtaining a feature extraction model to be trained and multiple pieces of sample information to be trained, the multiple pieces of sample information to be trained having respectively corresponding sample representation information, the multiple pieces of sample representation information having respectively corresponding information features, the target sample representation information corresponding to the target sample information to be trained being used to represent the target sample information to be trained, the target sample representation information corresponding to a third information feature, and the feature extraction model to be trained being used to extract the information features corresponding to the information;
[0182] Extracting the to-be-determined information features corresponding to the target sample information to be trained through the feature extraction model to be trained;
[0183] Adjusting the feature extraction model to be trained according to the similarity between the to-be-determined information features and the information features respectively corresponding to the multiple pieces of sample representation information to obtain the initial feature extraction model, and among the information features respectively corresponding to the multiple pieces of sample representation information, the similarity between the to-be-determined information features determined by the initial feature extraction model and the third information feature is the largest.
[0184] In a possible implementation, the sample information to be trained is image information, the sample representation information is text information, and the text information is used to express the information expressed by the corresponding image information.
[0185] In a possible implementation, the adjustment unit 903 is specifically configured to:
[0186] Taking the multiple information dimensions as the target information dimensions respectively, and extracting the first information features according to the target sample information through the target feature extraction model;
[0187] Determining the second information processing result corresponding to the first information features through the target feature processing model;
[0188] Adjusting the model parameters corresponding to the target feature extraction model according to the difference between the second information processing result and the target sample information processing result to obtain a feature extraction model.
[0189] In a possible implementation, the sample information is image information, and the information processing results respectively corresponding to the multiple information dimensions include any combination of an object contour recognition result, an image semantic recognition result, an object category recognition result, and an image character recognition result.
[0190] In a possible implementation, the apparatus further includes a third acquisition unit and a determination unit:
[0191] The third acquisition unit is configured to acquire a verification information set and a feature processing model respectively corresponding to multiple verification information dimensions. The target verification information set corresponding to the target verification information dimension includes multiple verification information. The multiple verification information has respectively corresponding sample information processing results. The second sample information processing result corresponding to the target verification information is determined based on the information expressed by the target verification information in the target verification information dimension. The multiple verification information dimensions do not include the multiple information dimensions. The feature processing model corresponding to the target verification information dimension is configured to determine an information processing result according to the information expressed by the information feature in the target verification information dimension. The target verification information dimension is any one of the multiple verification information dimensions;
[0192] The determination unit is configured to verify the feature extraction model through the verification information sets and the feature processing models respectively corresponding to the multiple verification information dimensions, and determine the versatility parameter corresponding to the feature extraction model. The versatility parameter is used to characterize the accuracy of the information feature extracted by the feature extraction model in the information expressed by the multiple verification information dimensions. The accuracy characterized by the versatility parameter is inversely correlated with the difference between the third information processing result and the second sample information processing result. The third information processing result is the information processing result determined by the feature processing model corresponding to the target verification information dimension according to the fourth information feature. The fourth information feature is the information feature extracted by the feature extraction model according to the target verification information.
[0193] An embodiment of the present application further provides a computer device. Please refer to Figure 10 As shown, the computer device may be a terminal device. Taking the terminal device as a mobile phone as an example:
[0194] Figure 10 Shown is a block diagram of a part of the structure of a mobile phone related to the terminal device provided by the embodiment of the present application. Refer to Figure 10, the mobile phone includes components such as a Radio Frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a Wireless Fidelity (WiFi) module 770, a processor 780, and a power supply 790. Those skilled in the art can understand that Figure 10 the mobile phone structure shown in
[0195] does not limit the mobile phone, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 10 The following specifically introduces each component of the mobile phone:
[0196] The RF circuit 710 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 780 for processing; in addition, the designed uplink data is sent to the base station. Usually, the RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0197] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 720 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0198] The input unit 730 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 731), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive and execute the commands sent by the processor 780. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 may also include other input devices 732. Specifically, the other input devices 732 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0199] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it is transmitted to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides a corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 10 the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.
[0200] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.
[0201] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and then converted into audio data. After the audio data is output and processed by the processor 780, it is sent to another mobile phone through the RF circuit 710, or the audio data is output to the memory 720 for further processing.
[0202] WiFi belongs to short-distance wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 770. It provides users with wireless broadband Internet access. Although Figure 10The WiFi module 770 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention as needed.
[0203] The processor 780 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720, it executes various functions of the mobile phone and processes data, thereby performing an overall detection of the mobile phone. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 780 either.
[0204] The mobile phone also includes a power supply 790 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 780 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0205] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0206] In this embodiment, the processor 780 included in the terminal device further has the following functions:
[0207] Obtain a target feature extraction model and feature processing models corresponding to multiple information dimensions. The target feature extraction model is used to extract information features corresponding to information, and the target feature processing model corresponding to the target information dimension is used to determine an information processing result according to the information expressed by the information features in the target information dimension. The information features are used to express the information expressed by the corresponding information in each information dimension;
[0208] Obtain sample information sets corresponding to the multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has respectively corresponding sample information processing results. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information is any one sample information in the target sample information set;
[0209] Taking the multiple information dimensions as the target information dimensions respectively, adjusting the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set, to obtain a feature extraction model, where the difference between the information processing result determined by the target feature processing model based on the first information feature and the target sample information processing result is less than a first preset threshold, and the first information feature is the information feature extracted by the feature extraction model from the target sample information.
[0210] The embodiments of the present application also provide a server. Please refer to Figure 11 as shown in Figure 11 which is a structural diagram of the server 800 provided by the embodiments of the present application. The server 800 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 822 (for example, one or more processors) and a memory 832, and one or more storage media 830 (for example, one or more mass storage devices) for storing application programs 842 or data 844. Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 822 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.
[0211] The server 800 may further include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0212] The steps performed by the server in the above embodiments may be based on Figure 11 the server structure shown in
[0213] The embodiments of the present application also provide a computer-readable storage medium for storing a computer program, and the computer program is used to execute any one of the model training methods described in the foregoing embodiments.
[0214] The embodiments of the present application also provide a computer program product including a computer program. When it runs on a computer device, it causes the computer device to execute the model training method described in any one of the above embodiments.
[0215] It can be understood that in the specific implementation manners of the present application, data related to user information (such as training information related to users) is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0216] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (abbreviation: ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.
[0217] It should be noted that the various embodiments in this specification are described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0218] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A model training method, characterized in that, The method includes: Obtaining a target feature extraction model and feature processing models corresponding to multiple information dimensions respectively. The target feature extraction model is used to extract information features corresponding to information. The target feature processing model corresponding to the target information dimension is used to determine an information processing result according to the information expressed by the information features in the target information dimension. The information features are used to express the information expressed by the corresponding information in each information dimension; Obtaining sample information sets corresponding to the multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has respectively corresponding sample information processing results. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information is any one of the sample information in the target sample information set; Taking the multiple information dimensions as the target information dimension respectively, adjusting the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model. The difference between the information processing result determined by the target feature processing model according to the first information feature and the target sample information processing result is less than a first preset threshold. The first information feature is the information feature extracted by the feature extraction model according to the target sample information.
2. The method according to claim 1, characterized in that, The target feature processing model is obtained through the following method: Obtaining an initial feature processing model corresponding to the target information dimension; Extracting a second information feature corresponding to the target sample information through an initial feature extraction model. The initial feature extraction model is used to extract information features corresponding to information; Determining a first information processing result corresponding to the second information feature through the initial feature processing model; Adjusting the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result to obtain the target feature processing model.
3. The method according to claim 2, wherein The adjusting the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result to obtain the target feature processing model includes: Adjusting the model parameters corresponding to the initial feature processing model according to the difference between the first information processing result and the target sample information processing result, so that the difference between the first information processing result and the target sample information processing result is less than a second preset threshold to obtain the target feature processing model. The second preset threshold is greater than the first preset threshold.
4. The method according to claim 3, wherein The initial feature extraction model is the target feature extraction model.
5. The method according to claim 3, wherein The target feature extraction model is generated by adding adjustable parameters to the initial feature extraction model. The adjustable parameters are used to participate in extracting the information features corresponding to the information. The number of parameters of the adjustable parameters is less than the number of parameters included in the initial feature extraction model. Taking the multiple information dimensions as the target information dimensions respectively, adjusting the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model, including: Taking the multiple information dimensions as the target information dimensions respectively, adjusting the adjustable parameters in the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model.
6. The method according to claim 2, wherein The adjustment of the model parameters of the initial feature processing models corresponding to the multiple information dimensions is performed in parallel.
7. The method according to claim 2, wherein The initial feature extraction model is obtained through the following training method: Obtaining a feature extraction model to be trained and multiple pieces of training sample information. The multiple pieces of training sample information have respectively corresponding sample representation information, and the multiple sample representation information have respectively corresponding information features. The target sample representation information corresponding to the target training sample information is used to represent the target training sample information. The target sample representation information corresponds to a third information feature. The feature extraction model to be trained is used to extract the information features corresponding to the information; Extracting the to-be-determined information features corresponding to the target training sample information through the feature extraction model to be trained; Adjusting the feature extraction model to be trained according to the similarity between the to-be-determined information features and the information features corresponding to the multiple sample representation information respectively, to obtain the initial feature extraction model. Among the information features corresponding to the multiple sample representation information respectively, the similarity between the to-be-determined information features determined by the initial feature extraction model and the third information feature is the largest.
8. The method according to claim 7, characterized in that The training sample information is image information, the sample representation information is text information, and the text information is used to express the information expressed by the corresponding image information.
9. The method according to claim 1, wherein Taking the multiple information dimensions as the target information dimensions respectively, adjusting the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model, including: Taking the multiple information dimensions as the target information dimensions respectively, and extracting the first information features according to the target sample information through the target feature extraction model; Determining the second information processing result corresponding to the first information feature through the target feature processing model; Adjusting the model parameters corresponding to the target feature extraction model according to the difference between the second information processing result and the target sample information processing result to obtain a feature extraction model.
10. The method according to claim 1, characterized in that, The sample information is image information, and the information processing results corresponding to the multiple information dimensions respectively include any combination of an object contour recognition result, an image semantic recognition result, an object category recognition result, and an image character recognition result.
11. The method according to claim 1, wherein The method further includes: Obtain the verification information sets and feature processing models corresponding to multiple verification information dimensions respectively. The target verification information set corresponding to the target verification information dimension includes multiple verification information, and the multiple verification information has respectively corresponding sample information processing results. The second sample information processing result corresponding to the target verification information is determined based on the information expressed by the target verification information in the target verification information dimension. The multiple verification information dimensions do not include the multiple information dimensions. The feature processing model corresponding to the target verification information dimension is used to determine the information processing result according to the information expressed by the information feature in the target verification information dimension. The target verification information dimension is any one of the multiple verification information dimensions. Verify the feature extraction model through the verification information sets and feature processing models corresponding to the multiple verification information dimensions respectively, and determine the generalizability parameter corresponding to the feature extraction model. The generalizability parameter is used to characterize the accuracy of the information feature extracted by the feature extraction model in the information expressed by the multiple verification information dimensions. The accuracy characterized by the generalizability parameter is inversely correlated with the difference between the third information processing result and the second sample information processing result. The third information processing result is the information processing result determined by the feature processing model corresponding to the target verification information dimension according to the fourth information feature, and the fourth information feature is the information feature extracted by the feature extraction model according to the target verification information.
12. A model training device, characterized in that, The device includes a first acquisition unit, a second acquisition unit, and an adjustment unit: The first acquisition unit is used to acquire the target feature extraction model and the feature processing models corresponding to multiple information dimensions respectively. The target feature extraction model is used to extract the information feature corresponding to the information. The target feature processing model corresponding to the target information dimension is used to determine the information processing result according to the information expressed by the information feature in the target information dimension. The information feature is used to express the information expressed by the corresponding information in each information dimension. The second acquisition unit is used to acquire the sample information sets corresponding to the multiple information dimensions respectively. The target sample information set corresponding to the target information dimension includes multiple sample information, and the multiple sample information has respectively corresponding sample information processing results. The target sample information processing result corresponding to the target sample information is determined based on the information expressed by the target sample information in the target information dimension. The target sample information is any one of the sample information in the target sample information set. The adjustment unit is used to use the multiple information dimensions as the target information dimension respectively, and adjust the model parameters corresponding to the target feature extraction model through the target feature processing model and the target sample information set to obtain a feature extraction model. The difference between the information processing result determined by the target feature processing model according to the first information feature and the target sample information processing result is less than the first preset threshold, and the first information feature is the information feature extracted by the feature extraction model according to the target sample information.
13. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is used to execute the model training method according to any one of claims 1-11 based on the instructions in the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the model training method according to any one of claims 1-11.
15. A computer program product including a computer program, when it runs on a computer device, causes the computer device to execute the model training method according to any one of claims 1-11.