VAE model training method, VAE model, data reconstruction method and electronic equipment
By selecting the structure of the encoder and decoder according to the downstream task type, the symmetry requirements of the VAE model are solved, and the problem of poor reconstruction performance of the existing VAE model is achieved, and better reconstruction performance and flexibility are achieved.
Patent Information
- Application Number
- CN202510153576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-27
AI Technical Summary
Due to the symmetry requirements, the network structure selection of the encoder and decoder is limited, resulting in poor reconstruction performance.
By selecting the appropriate encoder structure and decoder structure according to different downstream task types, an initial VAE model is built, and the initial VAE model is trained using the data to be trained to relieve the symmetry requirements.
It effectively improves the reconstruction performance of the VAE model, improves the information flow between the encoder and the decoder, and improves the flexibility of the model.
Smart Images

Figure CN120218155A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a VAE model training method, a VAE model, a data reconstruction method, and an electronic device. Background Art
[0002] In the field of artificial intelligence, data is an important core key. Generally, high-dimensional data in reality is compressed onto a low-dimensional feature layer for corresponding data processing, and then the low-dimensional feature data is output and converted according to the corresponding task requirements. In order to make more full use of the large amount of data generated in production, a symmetric VAE (Variational Auto Encoder) model is usually used for reconstruction. Generally, a VAE model consists of an encoder and a decoder. The encoder and the decoder use a symmetric structure to ensure that the data can be reconstructed to the original size and original dimension. However, the network structures suitable for the encoder and the decoder may be very different, and the symmetry in the symmetric VAE model limits the selection of the network structures of the encoder and the decoder, resulting in poor reconstruction performance of the model. Summary of the Invention
[0003] The present invention aims to at least partly solve one of the technical problems in the related art. To this end, the first object of the present invention is to propose a VAE model training method, which removes the symmetry requirement of the VAE model in the related art, and selects appropriate encoder structures and decoder structures according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0004] The second object of the present invention is to propose a VAE model.
[0005] The third object of the present invention is to propose a data reconstruction method.
[0006] The fourth object of the present invention is to propose a computer-readable storage medium.
[0007] The fifth object of the present invention is to propose an electronic device.
[0008] To achieve the above object, according to the first aspect embodiment of the present invention, a VAE model training method is proposed, including: obtaining training data to be trained and a downstream task type; respectively determining the structure of an encoder and the structure of a decoder according to the downstream task type, and constructing an initial VAE model according to the structure of the encoder and the structure of the decoder; and training the initial VAE model with the training data to be trained to obtain a VAE model.
[0009] According to the VAE model training method of an embodiment of the present invention, obtain the data to be trained and the downstream task type, and respectively determine the structure of the encoder and the structure of the decoder according to the downstream task type, and construct an initial VAE model according to the structure of the encoder and the structure of the decoder, and use the data to be trained to train the initial VAE model to obtain a VAE model. Thus, the symmetry requirement of the VAE model in the related art is removed, and a suitable encoder structure and decoder structure are selected according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0010] According to an embodiment of the present invention, the downstream task type includes a first task type and a second task type, wherein the feature compression level of the first task type is greater than the data recovery level of the first task type, and the feature compression level of the second task type is less than or equal to the data recovery level of the second task type.
[0011] According to an embodiment of the present invention, when the task type is the first task type, the encoder is a first network, and the decoder is a second network, wherein the number of parameters of the first network is greater than the number of parameters of the second network.
[0012] According to an embodiment of the present invention, the first network is a backbone network, and the second network is one of MAE and LinearProjection.
[0013] According to an embodiment of the present invention, when the task type is the second task type, the encoder is a third network, and the decoder is a fourth network, wherein the number of parameters of the fourth network is greater than the number of parameters of the third network.
[0014] According to an embodiment of the present invention, the third network is ResNet, and the fourth network adopts Unet.
[0015] According to an embodiment of the present invention, using the data to be trained to train the VAE model to obtain a VAE model includes: using the data to be trained to train the initial VAE model to obtain a training result; adjusting the model parameters of the encoder according to the training result until the accuracy of the encoder meets the first preset accuracy; adjusting the model parameters of the decoder according to the training result until the accuracy of the encoder meets the second preset accuracy to obtain a VAE model, wherein the first preset accuracy and the second preset accuracy are determined according to the downstream task type.
[0016] According to an embodiment of the present invention, the first preset accuracy when the downstream task type is the first task type is greater than the first preset accuracy when the downstream task type is the second task type, and the second preset accuracy when the downstream task type is the first task type is less than the second preset accuracy when the downstream task type is the second task type.
[0017] To achieve the above object, according to an embodiment of the second aspect of the present invention, a VAE model is proposed, including: an encoder and a decoder, wherein the architectures of the encoder and the decoder are determined based on the type of downstream task.
[0018] The VAE model according to an embodiment of the present invention includes an encoder and a decoder, wherein the architectures of the encoder and the decoder are determined based on the type of downstream task. Thus, the symmetry requirement of the VAE model in the related art is removed, and a suitable encoder structure and decoder structure are selected according to different types of downstream tasks, thereby effectively improving the reconstruction performance of the VAE model.
[0019] To achieve the above object, according to an embodiment of the third aspect of the present invention, a data reconstruction method is proposed, including: obtaining input data and the type of downstream task; obtaining a pre-trained VAE model according to the type of downstream task, wherein the architectures of the VAE models corresponding to different types of downstream tasks are different; using the VAE model to reconstruct the input data and output the reconstructed data.
[0020] The data reconstruction method according to an embodiment of the present invention obtains input data and the type of downstream task, obtains a pre-trained VAE model according to the type of downstream task, wherein the architectures of the VAE models corresponding to different types of downstream tasks are different, and uses the VAE model to reconstruct the input data and output the reconstructed data. Thus, the pre-trained VAE model removes the symmetry requirement of the VAE model in the related art, and a suitable encoder structure and decoder structure are selected according to different types of downstream tasks. Therefore, the reconstruction performance of the VAE model is better, and thus better reconstructed data can be obtained by using the VAE model.
[0021] To achieve the above object, according to an embodiment of the fourth aspect of the present invention, a computer-readable storage medium is proposed, on which a computer program is stored. When the computer program is processed by a processor, it executes the VAE model training method of any of the foregoing embodiments or the foregoing data reconstruction method.
[0022] The computer-readable storage medium according to an embodiment of the present invention, by executing the computer program of the above VAE model training method or the above data reconstruction method, removes the symmetry requirement of the VAE model in the related art, and selects a suitable encoder structure and decoder structure according to different types of downstream tasks, thereby effectively improving the reconstruction performance of the VAE model.
[0023] To achieve the above object, according to an embodiment of the fifth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a VAE model training program or a data reconstruction program stored in the memory and executable on the processor. When the processor executes the VAE model training program or the data reconstruction program, the VAE model training method of any of the foregoing embodiments or the foregoing data reconstruction method is implemented.
[0024] For the electronic device according to an embodiment of the present invention, by the processor executing the computer program of the above VAE model training method or the above data reconstruction method, the symmetry requirement of the VAE model in the related art is lifted, and a suitable encoder structure and decoder structure are selected according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0025] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic structural diagram of a symmetric VAE model in the related art;
[0027] Figure 2 is a schematic flowchart of a VAE model training method according to an embodiment of the present invention;
[0028] Figure 3 is a schematic structural diagram of a VAE model according to an embodiment of the present invention;
[0029] Figure 4 is a schematic flowchart of a data reconstruction method according to an embodiment of the present invention;
[0030] Figure 5 is a schematic system diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0032] It should be noted that this application is made by the inventor's understanding and research on the following problems:
[0033] In the field of artificial intelligence, data is a crucial core. High-dimensional data includes main features and redundant information, and the redundant information will increase the computational load of downstream tasks, resulting in a decrease in the efficiency of downstream tasks. Therefore, it is necessary to remove the redundant information in high-dimensional data. Thus, in related technologies, high-dimensional data is compressed onto a low-dimensional feature layer for corresponding data processing, and then the low-dimensional feature data is output and converted according to the corresponding task requirements. In order to make more full use of the large amount of data generated in production, a common solution is to use a self-supervised method to train a model that can perform feature extraction and feature processing on the data. The reconstruction task is a task based on the self-supervised method and is reconstructed based on a symmetric VAE model. Generally, the VAE model consists of an encoder and a decoder. The encoder usually includes convolutional layers and downsampling layers. The encoder is responsible for compressing and processing the data, mapping the data into the latent space, and the features follow a Gaussian distribution. The decoder usually includes transposed convolutional layers, upsampling layers, etc. The decoder is responsible for restoring the features mapped into the latent space that follows a Gaussian distribution to the original space. As Figure 1 shown, the encoder maps the input image into the latent space to obtain latent space features, and the latent space features follow a Gaussian distribution. The decoder restores the latent space features to the distribution of the original data to obtain a reconstructed image. In order to ensure that the data can be reconstructed to the original size and original dimension, the encoder and decoder use a symmetric structure, and the capabilities of the encoder and decoder should be approximated, that is, the number of parameters should be the same.
[0034] However, there are the following problems in the reconstruction based on the symmetric VAE model:
[0035] 1. Poor model flexibility: Due to the symmetric structure, when modifying the model structure, it is necessary to synchronously modify the encoder and decoder.
[0036] 2. Unbalanced information flow: In the symmetric structure, the number of network layers and parameters of the encoder and decoder are usually equal, which may lead to unbalanced information flow in both directions, resulting in the data reconstructed by the model tending to be smoothed data, seriously affecting the performance of the model.
[0037] 3. Limited model performance: In actual applications, the suitable network structures for the encoder and decoder may be different, and the symmetry in the symmetric VAE model limits the selection of network structures for the encoder and decoder, resulting in poor reconstruction performance of the model.
[0038] Based on this, the embodiments of the present invention provide a VAE model training method, a VAE model, a data reconstruction method, an electronic device, and a storage medium, which relieve the symmetry requirement of the VAE model in related technologies, and select appropriate encoder structures and decoder structures according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0039] The following describes a VAE model training method, a VAE model, a data reconstruction method, an electronic device, and a storage medium according to embodiments of the present invention with reference to the accompanying drawings.
[0040] Figure 2 It is a schematic flowchart of a VAE model training method according to an embodiment of the present invention. As Figure 2 shown, the VAE model training method includes:
[0041] S101, obtaining the data to be trained and the type of downstream task.
[0042] Specifically, since the VAE model is used to reconstruct data, the VAE model can adopt a self-supervised learning method, that is, the training data and labels are the same. Therefore, when training VAE models with different architectures, the same data to be trained can be used. The type of downstream task is determined according to the requirements of the downstream task for the reconstructed data, and the user can input the type of downstream task according to the downstream task.
[0043] In some embodiments, the type of downstream task includes a first task type and a second task type, where the feature compression level of the first task type is greater than the data recovery level of the first task type, and the feature compression level of the second task type is less than or equal to the data recovery level of the second task type.
[0044] Specifically, the feature compression level is used to characterize the performance requirements for the encoder, and the data recovery level is used to characterize the performance requirements for the decoder. The feature compression level of the first task type is greater than the data recovery level of the first task type, that is, the first task type has higher performance requirements for the encoder and lower performance requirements for the decoder; the feature compression level of the second task type is less than or equal to the data recovery level of the second task type, that is, the second task type has higher performance requirements for the decoder and lower performance requirements for the encoder.
[0045] For example, assume that the downstream task is a classification task. The classification task requires the reconstructed data to retain the main features and does not need to retain all features. Therefore, the requirement for the feature compression level is high, and the requirement for the data recovery level is low. So the type of downstream task is the first task type; assume that the downstream task is visualization. Visualization requires the reconstructed data to retain more features so that the visualized picture can be as consistent as possible with the original input picture. Therefore, the requirement for the feature compression level is low, and the requirement for the data recovery level is high. So the type of downstream task is the second task type.
[0046] In an alternative implementation, the type of downstream task can also be obtained by classifying the downstream task. The user can input the downstream task, and the downstream task is classified by a classification model to obtain the type of downstream task.
[0047] S102. Determine the structures of the encoder and the decoder respectively according to the downstream task type, and construct an initial VAE model based on the structures of the encoder and the decoder.
[0048] Specifically, different downstream task types have different performance requirements for the encoder and the decoder. If the performance requirement for the encoder is high, an encoder with a more complex structure can be selected. For example, an encoder with more network layers and more parameters can be selected. If the performance requirement for the encoder is low, an encoder with a simpler structure can be selected. For example, an encoder with fewer network layers and fewer parameters can be selected. If the performance requirement for the decoder is high, a decoder with a more complex structure can be selected. For example, a decoder with more network layers and more parameters can be selected. If the performance requirement for the decoder is low, a decoder with a simpler structure can be selected. For example, a decoder with fewer network layers and fewer parameters can be selected. Therefore, a suitable encoder structure and decoder structure can be selected according to the downstream task type. Then, generate the corresponding encoder and decoder according to the structures of the encoder and the decoder, and combine the encoder and the decoder to obtain the initial VAE model.
[0049] In some embodiments, when the task type is the first task type, the encoder is the first network and the decoder is the second network, where the number of parameters of the first network is greater than that of the second network.
[0050] It can be understood that the first task type has a higher performance requirement for the encoder and a lower performance requirement for the decoder. Therefore, the first network with more parameters needs to be selected for the encoder, and the second network with fewer parameters needs to be selected for the decoder.
[0051] Furthermore, in some embodiments, the first network is the backbone network, and the second network is one of MAE (Masked Auto Encoders) and Linear Projection.
[0052] It should be noted that the backbone network can be one of ViT (Vision Transformer, a neural network architecture based on self-attention mechanism), Swin Transformer (a vision model based on Transformer), and Flatten Transformer (a Transformer model with focused linear attention mechanism). The backbone network can also be other general networks, which are not restricted here specifically. Linear Projection can be the Linear Projection network in the SimMIM (Simplified Masked Image Modeling, a simple framework for masked image modeling) network. If the performance requirements for the decoder are very low, a one-layer Linear Projection can even be selected.
[0053] In some embodiments, when the task type is the second task type, the encoder is the third network and the decoder is the fourth network, where the number of parameters of the fourth network is greater than that of the third network.
[0054] Similarly, the second task type has lower requirements for the performance of the encoder and higher requirements for the performance of the decoder. Therefore, a third network with fewer parameters needs to be selected for the encoder, and a fourth network with more parameters needs to be selected for the decoder.
[0055] In some embodiments, the third network is ResNet (Residual Network), and the fourth network adopts Unet (a segmentation model).
[0056] It should be noted that the third network can be ResNet18, and the fourth network can be UnetPP (an improved Unet model).
[0057] S103: Train the initial VAE model using the data to be trained to obtain the VAE model.
[0058] Specifically, the initial VAE model is a general basic model, so its performance is poor. Therefore, it is necessary to use the data to be trained to adjust the encoder structure and the decoder structure respectively to obtain a VAE model with better performance. Since the VAE model in this embodiment removes the restriction of symmetry requirements, the encoder structure and the decoder structure can be adjusted differently.
[0059] In the above embodiments, the optimal encoder structure and decoder structure are selected according to the downstream task type. Therefore, the encoder and decoder do not need to adopt the same structure, which lifts the restriction of the symmetry requirement. The asymmetric structure can effectively improve the information flow in the encoder and decoder, thereby enhancing the flexibility and reconstruction performance of the VAE model.
[0060] In some embodiments, the VAE model is trained using the data to be trained to obtain the VAE model, including: training the initial VAE model using the data to be trained to obtain a training result; adjusting the model parameters of the encoder according to the training result until the accuracy of the encoder meets the first preset accuracy; adjusting the model parameters of the decoder according to the training result until the accuracy of the encoder meets the second preset accuracy to obtain the VAE model, where the first preset accuracy and the second preset accuracy are determined according to the downstream task type.
[0061] Specifically, training the VAE model using the data to be trained can obtain a training result, and the training result can be a loss function value, such as the reconstruction error (measuring the difference between the reconstructed image and the original image) and the KL divergence (Kullback-Leibler Divergence, relative entropy, used to measure the difference between the latent distribution and the standard normal distribution). Then, the model parameters of the encoder are first adjusted using the training result. After the accuracy of the encoder is adjusted to the first preset accuracy, the model parameters of the decoder are adjusted because the latent space features output by the encoder will affect the accuracy of the decoder. Therefore, adjusting the model parameters of the encoder first, and the model parameters can be adjusted for both the encoder and the decoder, where the model parameters include network parameters and the number of network layers, etc. Different downstream task types have different performance requirements for the encoder and decoder. For example, if it is the first task type, the accuracy requirement for the encoder is higher and the accuracy requirement for the decoder is lower; if it is the second task type, the accuracy requirement for the decoder is higher and the accuracy requirement for the encoder is lower. Therefore, the first preset accuracy and the second preset accuracy are determined according to the downstream task type.
[0062] In some embodiments, the first preset accuracy when the downstream task type is the first task type is greater than the first preset accuracy when the downstream task type is the second task type, and the second preset accuracy when the downstream task type is the first task type is less than the second preset accuracy when the downstream task type is the second task type.
[0063] It can be understood that the performance requirements of the first task type for the encoder are higher than those of the second task type for the encoder. Therefore, the first preset accuracy when the downstream task type is the first task type is greater than the first preset accuracy when the downstream task type is the second task type; the performance requirements of the second task type for the decoder are higher than those of the first task type for the decoder. Therefore, the second preset accuracy when the downstream task type is the first task type is less than the second preset accuracy when the downstream task type is the second task type.
[0064] In the above embodiments, different task types have different performance requirements for the encoder and the decoder. Therefore, when training the VAE model, different task types also have different accuracy requirements for the encoder and the decoder. Therefore, adjusting the accuracy of the encoder and the decoder according to the downstream task type can further improve the reconstruction performance of the VAE model.
[0065] In summary, according to the VAE model training method of the embodiments of the present invention, obtain the data to be trained and the downstream task type, and respectively determine the structure of the encoder and the structure of the decoder according to the downstream task type, and construct an initial VAE model according to the structure of the encoder and the structure of the decoder, and use the data to be trained to train the initial VAE model to obtain the VAE model. Thus, by selecting appropriate encoder structures and decoder structures according to different downstream task types, the symmetry requirements of the VAE model in the related art are removed, and the asymmetric structure can effectively improve the information flow in the encoder and the decoder, thereby improving the flexibility and reconstruction performance of the VAE model.
[0066] Corresponding to the above embodiments, an embodiment of the present invention further provides a VAE model. As Figure 3 shown, the VAE model 100 includes: an encoder 10 and a decoder 20, wherein the architectures of the encoder 10 and the decoder 20 are determined based on the downstream task type.
[0067] In some embodiments, the downstream task type includes a first task type and a second task type, wherein the feature compression level of the first task type is greater than the data recovery level of the first task type, and the feature compression level of the second task type is less than or equal to the data recovery level of the second task type.
[0068] In some embodiments, when the task type is the first task type, the encoder 10 is the first network, and the decoder 20 is the second network, wherein the number of parameters of the first network is greater than the number of parameters of the second network.
[0069] In some embodiments, the first network is the backbone network, and the second network is one of MAE and Linear Projection.
[0070] In some embodiments, when the task type is the second task type, the encoder 10 is the third network, and the decoder 20 is the fourth network, where the number of parameters of the fourth network is greater than that of the third network.
[0071] In some embodiments, the third network is ResNet, and the fourth network adopts Unet.
[0072] It should be noted that the specific implementation manners of the VAE model in the embodiments of the present invention correspond one by one to the specific implementation manners of the VAE model construction method in the foregoing embodiments of the present invention, and will not be elaborated herein.
[0073] The VAE model according to the embodiments of the present invention includes an encoder and a decoder, where the architectures of the encoder and the decoder are determined based on the downstream task type. Thus, the VAE model eliminates the symmetry requirement of the VAE model in the related art, and selects appropriate encoder structures and decoder structures according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0074] Corresponding to the above embodiments, the embodiments of the present invention further provide a data reconstruction method. As Figure 4 shown, the data reconstruction method includes:
[0075] S201, obtaining input data and a downstream task type.
[0076] S202, obtaining a pre-trained VAE model according to the downstream task type, where the architectures of the VAE models corresponding to different downstream task types are different.
[0077] S203, reconstructing the input data by using the VAE model and outputting reconstructed data.
[0078] According to the data reconstruction method of the embodiments of the present invention, input data and a downstream task type are obtained, a pre-trained VAE model is obtained according to the downstream task type, where the architectures of the VAE models corresponding to different downstream task types are different, and the input data is reconstructed by using the VAE model and reconstructed data is output. Thus, the pre-trained VAE model eliminates the symmetry requirement of the VAE model in the related art, selects appropriate encoder structures and decoder structures according to different downstream task types. Therefore, the reconstruction performance of the VAE model is good, and thus good reconstructed data can be obtained by using the VAE model.
[0079] Corresponding to the above embodiments, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is processed by a processor, it executes the VAE model training method or the foregoing data reconstruction method in any of the foregoing embodiments.
[0080] According to the computer-readable storage medium of an embodiment of the present invention, by executing the computer program of the above VAE model training method or the above data reconstruction method, the symmetry requirement of the VAE model in the related art is lifted, and a suitable encoder structure and decoder structure are selected according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0081] Corresponding to the above embodiment, an embodiment of the present invention further provides an electronic device. As Figure 5 shown, the electronic device 200 includes a memory 210, a processor 220, and a VAE model training program or a data reconstruction program stored on the memory 210 and executable on the processor 220. When the processor 220 executes the VAE model training program or the data reconstruction program, the VAE model training method of any of the foregoing embodiments or the foregoing data reconstruction method is implemented.
[0082] According to the electronic device of an embodiment of the present invention, by the processor executing the computer program of the above VAE model training method or the above data reconstruction method, the symmetry requirement of the VAE model in the related art is lifted, and a suitable encoder structure and decoder structure are selected according to different downstream task types, thereby effectively improving the reconstruction performance of the VAE model.
[0083] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0084] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0085] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0086] In addition, the terms such as "first" and "second" used in the embodiments of the present invention are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the technical features indicated in this embodiment. Thus, the features defined with terms such as "first" and "second" in the embodiments of the present invention can explicitly or implicitly indicate that at least one such feature is included in this embodiment. In the description of the present invention, the meaning of the word "plurality" is at least two or more than two, such as two, three, four, etc., unless otherwise specifically defined in the embodiment.
[0087] In the present invention, unless otherwise clearly specified or limited in the embodiments, the terms such as "installed", "connected", "connected", and "fixed" appearing in the embodiments should be understood in a broad sense. For example, the connection can be a fixed connection, a detachable connection, or integrated. It can be understood that it can also be a mechanical connection, an electrical connection, etc.; of course, it can also be directly connected, or indirectly connected through an intermediate medium, or it can be the communication inside two elements, or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific implementation situations.
[0088] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as a limitation to the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A VAE model training method, characterized in that: include: Obtain the data to be trained and the downstream task type; Determine the structure of the encoder and the structure of the decoder respectively according to the downstream task type, and construct an initial VAE model according to the structure of the encoder and the structure of the decoder; The initial VAE model is trained using the data to be trained to obtain a VAE model.
2. The method according to claim 1, characterized in that The downstream task type includes a first task type and a second task type, wherein the feature compression level of the first task type is greater than the data recovery level of the first task type, and the feature compression level of the second task type is less than or equal to the data recovery level of the second task type.
3. The method according to claim 2, characterized in that When the task type is the first task type, the encoder is a first network, and the decoder is a second network, wherein the parameter amount of the first network is greater than the parameter amount of the second network.
4. The method according to claim 3, characterized in that The first network is a backbone network, and the second network is one of MAE and Linear Projection.
5. The method according to claim 2, characterized in that: When the task type is the second task type, the encoder is a third network, and the decoder is a fourth network, wherein the parameter amount of the fourth network is greater than the parameter amount of the third network.
6. The method according to claim 5, characterized in that The third network is ResNet, and the fourth network is Unet.
7. The method according to claim 2, characterized in that The VAE model is trained using the data to be trained to obtain a VAE model, including: Using the data to be trained to train the initial VAE model to obtain a training result; Adjusting the model parameters of the encoder according to the training result until the accuracy of the encoder meets the first preset accuracy; The model parameters of the decoder are adjusted according to the training results until the accuracy of the encoder meets the second preset accuracy to obtain the VAE model, wherein the first preset accuracy and the second preset accuracy are determined according to the downstream task type.
8. The method according to claim 7, characterized in that The first preset accuracy when the downstream task type is the first task type is greater than the first preset accuracy when the downstream task type is the second task type, and the second preset accuracy when the downstream task type is the first task type is less than the second preset accuracy when the downstream task type is the second task type.
9. A VAE model, characterized in that: include: An encoder and a decoder, wherein the architecture of the encoder and the architecture of the decoder are determined based on a downstream task type.
10. A data reconstruction method, characterized in that: include: Get input data and downstream task types; Obtain a pre-trained VAE model according to the downstream task type, wherein different downstream task types correspond to different VAE model architectures; The VAE model is used to reconstruct the input data and output the reconstructed data.
11. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is processed by a processor, the VAE model training method described in any one of claims 1 to 8 or the data reconstruction method described in claim 10 is executed.
12. An electronic device, characterized in that: It includes a memory, a processor, and a VAE model training program or a data reconstruction program stored in the memory and executable on the processor. When the processor executes the VAE model training program or the data reconstruction program, it implements the VAE model training method described in any one of claims 1 to 8 or the data reconstruction method described in claim 10.