Image generation model training method, audio-based image generation method, and device
By adjusting the image generation model using age and gender labels from the training data, the generated images are made to match the characteristics of the audio, solving the problem of the lack of semantic connection between audio processing and image generation in existing technologies, improving user experience and reducing annotation costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
- Filing Date
- 2024-08-02
- Publication Date
- 2026-07-24
Smart Images

Figure CN118861353B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing, and more particularly to image generation model training methods, audio-based image generation methods and devices. Background Technology
[0002] To provide users with a better user experience, existing audio applications typically offer a song profile generation function, which allows users to generate a user profile for songs in their playlists.
[0003] Generally, in existing technologies, audio processing and image generation are usually performed separately. For example, audio features are extracted separately and a user avatar is generated based on these features. This approach cannot determine a clear semantic relationship between the generated avatar and the input audio features, or even any clear correlation at all. Furthermore, multimodal training data requires significant annotation costs. In other words, the avatars generated by existing technologies do not match the songs in a user's playlist well, resulting in a poor user experience when using the song profiling function. Summary of the Invention
[0004] This application provides an image generation model training method, an audio-based image generation method, and an apparatus for generating portraits that match the characteristics of human voices in corresponding image samples.
[0005] The first aspect of this application provides a method for training an image generation model, comprising:
[0006] Acquire training data, which includes first audio data, second audio data, image data, and age and gender labels corresponding to the first audio data and image data;
[0007] An audio classifier is trained using the age and gender labels in the training data and the first audio data.
[0008] An image classifier is trained using the age and gender labels in the training data and the image data.
[0009] The second audio data is processed based on a pre-trained image generation model to obtain a predicted image corresponding to the second audio data. The first audio data and the second audio data are not completely the same.
[0010] The second audio data is input into the audio classifier to obtain the predicted human voice feature label corresponding to the second audio data, and the predicted image is input into the image classifier to obtain the predicted human image feature label corresponding to the predicted image.
[0011] Based on the predicted human voice feature labels and the predicted human image feature labels, the parameters of the pre-trained image generation model are adjusted to obtain the trained image generation model.
[0012] In some specific implementations, adjusting the parameters of the pre-trained image generation model based on the predicted human voice feature labels and the predicted human image feature labels to obtain the trained image generation model includes:
[0013] The distance between the predicted human voice feature labels and the predicted human image feature labels is used as the first training loss, and the parameters of the pre-trained image generation model are adjusted based on the first training loss to obtain the trained image generation model.
[0014] or,
[0015] The predicted human voice feature labels and the predicted human image feature labels are input into the matching model to obtain the degree of matching between the predicted human voice feature labels and the predicted human image feature labels;
[0016] The matching degree between the predicted human voice feature labels and the predicted human image feature labels is used as the first training loss, and the parameters of the pre-trained image generation model are adjusted based on the first training loss to obtain the trained image generation model.
[0017] In some specific implementations, training an audio classifier using the age and gender labels from the training data and the first audio data includes:
[0018] Input the first audio data into the initial audio classifier to obtain the predicted age and gender labels corresponding to the first audio data;
[0019] The parameters of the initial audio classifier are adjusted based on the predicted age and gender labels corresponding to the first audio data and the distance between the age and gender labels in the training data to obtain the audio classifier;
[0020] The step of training an image classifier using the age and gender labels in the training data and the image data includes:
[0021] The image data is input into an initial image classifier to obtain the predicted age and gender labels corresponding to the image data;
[0022] The parameters of the initial image classifier are adjusted based on the predicted age and gender labels corresponding to the image data and the distance between the age and gender labels in the training data to obtain the image classifier.
[0023] In some specific implementations, obtaining training data includes:
[0024] Acquire multiple existing audio data sets;
[0025] A portion of the existing audio data is used as the first audio data, and image data corresponding to each first audio data input by the user is obtained, as well as age and gender labels corresponding to the first audio data and image data;
[0026] The existing audio data other than the first audio data is identified as the second audio data.
[0027] In some specific implementations, the method further includes:
[0028] The first audio data is input into the initial image generation model to obtain the predicted image corresponding to the first audio data.
[0029] The distance between the predicted image corresponding to the first audio data and the image data in the training data is used as the second training loss;
[0030] The parameters of the initial image generation model are adjusted based on the second training loss to obtain the pre-trained image generation model.
[0031] According to the method, the predicted human voice feature label includes age and gender labels, and the predicted human image feature label is consistent with the predicted human voice feature label.
[0032] A second aspect of this application provides an audio-based image generation method, including:
[0033] Acquire the audio data to be processed;
[0034] The audio data to be processed is input into a trained image generation model to obtain an image corresponding to the audio data to be processed. The trained image generation model is trained based on the training method of the image generation model.
[0035] In some specific implementations, the method further includes:
[0036] In response to a user-initiated instruction to generate a vocal image, at least one song is determined from the user's playlist and / or the songs sung by the user as the audio data to be processed.
[0037] A third aspect of this application provides a computer device, including:
[0038] Central processing unit, memory, and input / output interfaces;
[0039] The memory is either a short-term storage memory or a persistent storage memory;
[0040] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in the first or second aspect.
[0041] A fourth aspect of this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first or second aspect.
[0042] A fifth aspect of this application provides a computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the method as described in the first or second aspect.
[0043] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: Training data is acquired, including first audio data, second audio data, image data, and age and gender labels corresponding to the first audio data and image data; an audio classifier is trained using the age and gender labels in the training data and the first audio data; an image classifier is trained using the age and gender labels in the training data and the image data; the second audio data is processed based on a pre-trained image generation model to obtain a predicted image corresponding to the second audio data, where the first audio data and the second audio data are not entirely identical; the second audio data is input into the audio classifier to obtain predicted human voice feature labels corresponding to the second audio data, and the predicted image is input into the image classifier to obtain predicted human image feature labels corresponding to the predicted image; the parameters of the pre-trained image generation model are adjusted based on the predicted human voice feature labels and the predicted human image feature labels to obtain a trained image generation model. Thus, based on the trained image generation model, an image matching the corresponding human image features and the human voice features of the input audio data can be generated, meaning that there is sufficient semantic connection between the generated image and the input audio data, which can improve the user experience. Furthermore, in the training process of the pre-trained image generation model in this embodiment, it is not necessary to use pre-labeled age and gender tags or image data. This means that during the training of the pre-trained image generation model, unlabeled audio data can be used as the second audio data, which greatly reduces the labeling cost of training data and effectively improves the model training efficiency. Attached Figure Description
[0044] Figure 1 This is a framework diagram of a training method for an image generation model disclosed in an embodiment of this application;
[0045] Figure 2 This is a schematic flowchart illustrating a training method for an image generation model disclosed in an embodiment of this application.
[0046] Figure 3This is a schematic flowchart of an image generation method disclosed in an embodiment of this application;
[0047] Figure 4 This is a schematic diagram of the structure of a computer device disclosed in an embodiment of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0049] This application provides an image generation model training method, an audio-based image generation method, and an apparatus for generating portraits that match the characteristics of human voices in corresponding image samples.
[0050] Please see Figure 1 To better implement the image generation model training method of this application embodiment, this application embodiment provides an image generation model training framework. Figure 1 Taking the framework shown as an example, this application embodiment requires at least two stages of training on the initial image generation model in order to obtain a trained image generation model.
[0051] Phase 1: Train the initial image generation model using the second training loss to obtain a pre-trained image generation model. The second training loss is the consistency between the predicted image (output by the initial image generation model based on the input audio data) and the corresponding image data. For example, the consistency between the predicted image corresponding to the first audio data input into the initial image generation model and the image data in the training data can be used as the second training loss. It should be noted that in practical applications, in addition to using the first training data, different audio data that meets the training requirements but is different from the training image classifier and audio classifier can be used to train the initial image generation model; no specific limitations are made here. When the matching degree between the model's output image and the corresponding image data is considered sufficient, training ends and the image generation model is determined as the pre-trained image generation model.
[0052] The second stage involves training the pre-trained image generation model using the first training loss to obtain a trained image generation model. The first training loss represents the matching degree between the voice feature labels (output from the audio classifier after inputting the audio data) and the portrait feature labels (output from the image classifier after inputting the predicted image corresponding to the audio data). The predicted image corresponding to the audio data is the image output by the pre-trained image generation model after processing the aforementioned audio data. Similarly, when the matching degree between the portrait features in the predicted image output by the model and the voice features in the audio data is considered sufficient, training ends, and the image generation model is determined as the trained image generation model.
[0053] Please see Figure 2 Based on the aforementioned image generation model training framework, the image generation model training method of this application includes the following steps:
[0054] 201. Obtain training data, which includes first audio data, second audio data, image data, and age and gender labels corresponding to the first audio data and image data.
[0055] First, it should be noted that age and gender labels can be labels that directly reflect age and gender, such as: young woman, middle-aged woman, young man, and middle-aged man; or labels that indirectly reflect age and gender, such as: mature woman, sweet woman, mature man, and weathered man; or they can be composed of two labels: age label (including but not limited to teenager label, youth label, middle-aged label, and elderly label) and gender label (including but not limited to male label and female label).
[0056] As can be seen from the foregoing embodiments, this application actually performs two-stage training. Therefore, this application needs to obtain the following for the first stage of training: first audio data, image data, and age and gender labels corresponding to the first audio data and image data; while the second stage of training only requires the second audio data.
[0057] It should be noted that there is also a one-to-one correspondence between the first audio data and the image data. Generally, multiple first audio data are needed to train the image generation model in multiple rounds. Each first audio data, the image data that has a one-to-one correspondence with the first audio data, and the age and gender labels that also have a correspondence with the first audio data and the corresponding image data constitute a set of training samples, which can be used to train the model in one round.
[0058] Furthermore, the audio data referred to in this application (including but not limited to the first audio) can be the audio itself or the features of the audio itself, without limitation here. If the audio data is the audio itself, the audio data input to the audio classifier and the image generation model still needs to undergo an audio feature extraction step before it can be further processed by the network (in the model or classifier). The specific audio feature extraction network can be built into the audio classifier and the image generation model respectively; or, if the audio data is the features of the audio, the audio feature extraction step can be implemented by a separate neural network, without limitation here.
[0059] The audio features can be MFCC features, loudness, rhythm, and / or difference features, etc. If multiple features are included, they can be concatenated and / or averaged over time to obtain the (final) audio features. Alternatively, deep learning methods can be used to directly extract audio features (vectors). The original audio is processed through short-time frame segmentation, Fourier transform, deep learning, and / or feature learning to extract features (vectors) with high-level semantic features; specific limitations are not specified here.
[0060] Similarly, the image data referred to in this application includes, but is not limited to, the image itself or the features of the image. If the image data is the image itself, the image data input to the image classifier needs to go through the step of image feature extraction before it can be further processed by the network (in the model or classifier). The specific network for implementing image feature extraction can be built into the image classifier or implemented by an independent neural network, which is not limited here.
[0061] 202. Use the age and gender labels in the training data and the first audio data to train an audio classifier.
[0062] It is understood that during the training process of the audio classifier, this embodiment of the application uses the first audio data as input, and the training objective is to make the voice feature labels output by the audio classifier as close as possible to the age and gender labels corresponding to the input first audio data. The voice feature labels are trained using age and gender labels, and can also represent the age and gender of the input first audio data.
[0063] 203. Use the age and gender labels in the training data and the image data to train an image classifier.
[0064] Similarly, in the image classifier training process, this embodiment of the application uses image data as input, and the training objective is to make the facial feature labels output by the image classifier as close as possible to the age and gender labels corresponding to the input image data. The facial feature labels are trained using age and gender labels, and can also represent the age and gender of the person in the input image data.
[0065] 204. Based on a pre-trained image generation model, process the second audio data to obtain the predicted image corresponding to the second audio data. The first audio data and the second audio data are not completely the same.
[0066] As can be seen from the aforementioned image generation model training framework, the pre-trained image generation model can achieve the following: the predicted image generated based on the input audio data has a high degree of consistency with the image data corresponding to the audio data; however, the pre-trained image generation model cannot guarantee sufficient semantic relevance between the generated predicted image and the input audio data. Therefore, in this embodiment, the pre-trained image generation model needs to be further trained.
[0067] Specifically, the second audio data required for this stage of training is input into the pre-trained image generation model, and the pre-trained image generation model will output the predicted image, that is, the predicted image corresponding to the second audio data.
[0068] It's important to note that since the second audio data doesn't require corresponding image data or age and gender labels in the second training phase, unlabeled audio data can be used. This reduces the cost of building training data, improves training efficiency, and achieves the second training phase at extremely low cost. Furthermore, if needed, the first audio data used for training in the first phase can be used as the second audio data in the second training phase.
[0069] 205. Input the second audio data into the audio classifier to obtain the predicted human voice feature label corresponding to the second audio data, and input the predicted image into the image classifier to obtain the predicted human image feature label corresponding to the predicted image.
[0070] It is understandable that there are certain commonalities between the vocal and visual characteristics of the same person; in other words, some of a person's traits are reflected in both their appearance and voice. Based on these common traits, this application embodiment labels the visual and vocal characteristics, and calculates the degree of matching between them by extracting the vocal characteristics (labels) from the audio and the visual characteristics (labels) from the image data using a classifier. The aforementioned common characteristics may include, but are not limited to, age and gender.
[0071] Because of the training in step 202, this application embodiment already has the ability to accurately identify the vocal characteristics of the speaker or singer corresponding to the second audio data, which is reflected by predicting vocal characteristic labels; similarly, because of the training in step 203, this application embodiment already has the ability to accurately identify the portrait characteristics of the person in the predicted image corresponding to the second audio data, which is reflected by predicting portrait characteristic labels.
[0072] 206. Adjust the parameters of the pre-trained image generation model based on the predicted human voice feature labels and the predicted human image feature labels to obtain the trained image generation model.
[0073] It is understandable that there is a strong semantic correlation between a person's voice and their image; that is, there should be a strong consistency between the characteristics of a person's voice and the characteristics of their image. Therefore, this application uses the degree of matching between the predicted image feature labels of the person in the predicted portrait and the predicted voice feature labels in the audio data as an indicator to evaluate the quality of the image generation model.
[0074] The training objective of this step is to ensure that the predicted facial feature labels of people in the predicted images are consistent with the predicted voice feature labels in the audio data. Therefore, in this embodiment, the parameters of the pre-trained image generation model are adjusted based on the degree of matching between the predicted voice feature labels and the predicted facial feature labels until a well-trained image generation model is obtained.
[0075] The purpose of model training is to continuously approach the training objective and improve model quality. As described in the foregoing embodiments, the purpose of training the pre-trained image generation model in this application is to ensure that the predicted portrait output by the model appears to produce the human voice corresponding to the input audio data, or in other words, that the person appears to like the human voice corresponding to the input audio data. Therefore, the parameters of the pre-trained image generation model are adjusted based on the predicted human voice feature labels and the predicted human portrait feature labels, and a new round of training is performed until the convergence condition is met. At this point, the model training ends, and the image generation model at that time is determined as the trained image generation model. The convergence condition can be that the first loss value is less than a preset loss threshold and / or the number of training rounds reaches a preset round threshold, which can be configured as needed and is not limited here.
[0076] In addition, it is understood that each stage of model training requires multiple training samples for multiple rounds of iteration. Each training session in each training stage in this application embodiment can be performed in the manner described in the relevant embodiments of this application, and is not limited here.
[0077] In this embodiment, a pre-trained image generation model can generate images that match the facial features of the input audio data with the vocal features of the input audio data. This means that the generated images and the input audio data have sufficient semantic connection, improving user experience. Furthermore, during the training of the pre-trained image generation model, this embodiment does not require pre-labeled age and gender tags or image data. This means that unlabeled audio data can be used as secondary audio data during the training of the pre-trained image generation model, significantly reducing the labeling cost of training data and effectively improving model training efficiency. Generally, the method in this embodiment can be used to generate images, including but not limited to album covers and song portraits, that require a certain correlation with the audio data.
[0078] In some specific implementations, step 206 can be implemented by referring to the following steps: using the distance between the predicted voice feature labels and the predicted image feature labels as the first training loss, and adjusting the parameters of the pre-trained image generation model based on the first training loss to obtain a trained image generation model; or, inputting the predicted voice feature labels and the predicted image feature labels into the matching model to obtain the matching degree between the predicted voice feature labels and the predicted image feature labels; using the matching degree between the predicted voice feature labels and the predicted image feature labels as the first training loss, and adjusting the parameters of the pre-trained image generation model based on the first training loss to obtain a trained image generation model.
[0079] In simple terms, the degree of matching between predicted human voice feature labels and predicted human image feature labels can be determined by the distance between the two labels or by a matching model. The degree of matching between the two labels serves as the first training loss in one round of training. The trained image generation model can adjust its own model parameters using the first training loss to continuously reduce the first training loss and reach convergence. The distance between the two labels can be calculated using any distance algorithm such as cosine distance, Jaccard distance, or Euclidean distance, or any loss function that can calculate the loss of a classification problem, such as the cross-entropy loss function or the cosine similarity loss function. No limitation is imposed here.
[0080] Furthermore, step 202 can be implemented as follows: input the first audio data into an initial audio classifier to obtain the predicted age and gender labels corresponding to the first audio data; adjust the parameters of the initial audio classifier based on the predicted age and gender labels corresponding to the first audio data and the distance between the age and gender labels in the training data to obtain the audio classifier; similarly, step 203 can be implemented as follows: input the image data into an initial image classifier to obtain the predicted age and gender labels corresponding to the image data; adjust the parameters of the initial image classifier based on the predicted age and gender labels corresponding to the image data and the distance between the age and gender labels in the training data to obtain the image classifier.
[0081] Based on the above, the training of audio classifiers is similar to that of image classifiers. The training process for classifiers will be explained below.
[0082] Specifically, the audio classifier is continuously trained using the first audio data until a well-trained audio classifier is obtained; and the image classifier is continuously trained using image data until a well-trained image classifier is obtained. It should be noted that the training input for the audio classifier is the first audio data with corresponding age and gender labels, and the training data for the image classifier is image data with corresponding age and gender labels. The training inputs used by the two classifiers can have a one-to-one correspondence or not; this embodiment does not limit this.
[0083] Furthermore, image classifiers first need to extract image features from the input image (including but not limited to real portraits and predicted images), and then classify the image based on these features to obtain the characteristic labels of the people in the corresponding images. Generally, image feature extraction can be achieved using deep convolutional neural networks, including but not limited to VGGNet or ResNet.
[0084] It should also be noted that, when training the image classifier, in addition to using manually selected and labeled image data that has a one-to-one correspondence with the first audio data, a pre-trained image generation model can also be used to generate a predicted image based on any audio data, as long as the predicted image has corresponding age and gender labels (such as a predicted image generated by a pre-trained image generation model based on any first audio data). This application does not impose any limitations on this. Furthermore, the distance between two labels can be calculated using any distance algorithm such as cosine distance, Jaccard distance, or Euclidean distance, or it can be calculated using any loss function that can calculate the loss of a classification problem, such as the cross-entropy loss function or the cosine similarity loss function. This is not limited here.
[0085] In a specific embodiment, the loss function for calculating the loss of a classification problem is designed as follows:
[0086]
[0087] Where P = [P0,…P C-1 ], P i Let y represent the probability that a sample belongs to the i-th class. If the true class of the sample is the i-th class, then y i =1, otherwise y i =0.
[0088] Furthermore, the pre-trained image generation model can be implemented as follows: input the first audio data into the initial image generation model to obtain the predicted image corresponding to the first audio data; use the distance between the predicted image corresponding to the first audio data and the image data in the training data as the second training loss; adjust the parameters of the initial image generation model based on the training loss to obtain the pre-trained image generation model.
[0089] Specifically, the initial image generation model is continuously trained using the first audio data and corresponding image data until a well-trained image generation model is obtained. The calculation of the second training loss can refer to the aforementioned embodiments related to the calculation of the first training loss, and will not be repeated here. It should be noted that the initial image generation model in this application can be any deep learning network capable of generating images; there is no limitation here. However, if a Generative Adversarial Network (GAN) is used as the framework for the initial image generation model, the loss function of the generator G and discriminator D during the training process in the GAN network is as follows:
[0090]
[0091] Where V(D,G) represents the difference between the output predicted image and the corresponding image data of the input audio data, and the loss function is binary cross-entropy. Let G be a fixed discriminator, and D be a generator trained by maximizing the cross-entropy loss function V(D,G). The training objective of D is to correctly distinguish between the image data x corresponding to the input audio data and the predicted output image G(z). The stronger the discriminative power of D, the larger D(x) will be. Let D be a fixed discriminator, and G be the generator. The generator should minimize the cross-entropy loss V(D,G) between real and fake images while maximizing the discriminator's cross-entropy loss. The quality of the generated image is evaluated as closely as possible to the quality of the output predicted image to the image data corresponding to the input audio data. x~P data (x) represents the image data corresponding to the input audio data, z ~ P z(z) represents a sample generated from Gaussian-distributed noise splicing audio features.
[0092] Ultimately, the obtained generator is the image generation model pre-trained in this application.
[0093] It should be noted that, in addition to being trained in the manner described above, the pre-trained image generation model can also be an existing image generation model used as the pre-trained image generation model for this application; no limitation is made here.
[0094] Based on the foregoing embodiments, in some specific implementations, the embodiments of this application can construct training data in the following ways: acquire multiple existing audio data; use a portion of the existing audio data as first audio data, and acquire image data corresponding to each first audio data input by the user, as well as age and gender labels corresponding to the first audio data and image data; determine the existing audio data other than the first audio data as second audio data.
[0095] Specifically, a large amount of existing audio data can be acquired through open-source licensing or user authorization. Then, a portion of the existing audio data is selected as the first audio data based on requirements; for example, the number of existing audio data selected can be determined based on annotation capabilities. Next, image data input by the user for each piece of the first audio data, along with corresponding age and gender labels, is obtained through manual annotation. Finally, considering the isolation of training samples between different stages, any number of existing audio data that meet the requirements, excluding the first audio data, can be directly used as the second audio data; for example, the number of second audio data required for the second training stage can be determined based on the sample size required for the second training stage and / or the actual training effect of the first training stage.
[0096] The preceding text described various embodiments of the training method for the image generation model of this application. Based on the aforementioned training method for the image generation model, the following text provides an audio-based image generation method of this application, which includes the following steps: acquiring audio data to be processed; inputting the audio data to be processed into the trained image generation model to obtain the image corresponding to the audio data to be processed. The trained image generation model is trained based on any of the image generation model training methods in the aforementioned embodiments.
[0097] Since the image generation model trained based on the aforementioned embodiments can ensure that the generated image contains human portrait features matching the vocal characteristics reflected in the audio data, it meets the user's need to generate song portraits based on songs, or in other words, to generate album art for songs. Therefore, after obtaining the audio data to be processed, it can be input into the trained image generation model to obtain the image corresponding to the audio data to be processed. The key is that the portrait features of the people in the generated image match the vocal characteristics of the input audio data to be processed.
[0098] Specifically, the audio-based image generation process described in this application can be found in [reference needed]. Figure 3 , Figure 3 The image shown is the image generation model trained in this application. If there is a song called "Twilight" sung by a woman in her thirties, inputting that song into the trained image generation model can produce an image like this. Figure 3 The image shown in the lower right corner indicates that the person in the generated image is a woman in her thirties, ensuring a good match between the input audio data and the output image. Despite Figure 3 The image shown is in color, but in practical applications, if the image data used during the training phase is black and white, the trained image generation model can also output a black and white image.
[0099] The image generation model provided in this application, which can deeply integrate the characteristics of human voices in audio data, can be applied to movies and games to automatically design appearances based on the characteristics of human voices of characters, or to provide users with personalized music and video experiences, or even to create visual representations of audio content in virtual reality or augmented reality. It has great potential value in different fields.
[0100] In some specific implementations, users can initiate a vocal profile generation command, either directly through relevant pages or buttons, or indirectly by having the command automatically triggered after the user performs certain actions. Then, at least one song can be selected from the user's playlist and / or the songs the user has sung as the audio data to be processed. Specifically, if the user's playlist contains multiple songs, or the user has sung multiple songs, any song selected as needed, such as the user's favorite song or the song with the highest performance score, can be used as the audio data to be processed. This audio data is then input into a trained image generation model to generate its corresponding image, serving as the song profile. The user's favorite song can be the song with the highest recent listening frequency or the song with the highest total listening count; this can be configured as needed.
[0101] Building upon the aforementioned embodiments, to enhance the robustness of image generation and the diversity of generated images, random noise (such as...) can be added to the input of the image generation model during training. Figure 1 As shown in the diagram, random noise is added to the input of the trained image generation model during application to introduce minor perturbations. This ensures that even with the same audio data input, the output images will not be identical due to the random noise, enriching the user experience. Specifically, the proportion of random noise to the input of the image generation model should not exceed a preset interference threshold to avoid affecting the accuracy of image generation.
[0102] It is understood that the voice feature tags, image feature tags, predicted voice feature tags, and predicted image feature tags referred to in the foregoing embodiments of this application can be tags consistent with the age and gender tags included in the training data in step 201. For example, the voice feature tag and the predicted voice feature tag can be a tag that directly reflects age and gender, such as: young woman, middle-aged woman, young man, and middle-aged man; or they can be a tag that indirectly reflects age and gender, such as: mature woman, sweet woman, mature man, and weathered man; or they can be directly composed of two tags: age tag (specifically including but not limited to teenager tag, youth tag, middle-aged tag, and elderly tag) and gender tag (specifically including but not limited to male tag and female tag). Similarly, the portrait feature tag and the predicted portrait feature tag can be a tag that directly reflects age and gender, such as: young woman, middle-aged woman, young man, and middle-aged man; or it can be a tag that indirectly reflects age and gender, such as: mature woman, sweet woman, mature man, and weathered man; or it can be composed of two tags: age tag (including but not limited to teenager tag, youth tag, middle-aged tag, and elderly tag) and gender tag (including but not limited to male tag and female tag).
[0103] It should be noted that when the target is a user, the audio data and playlists and other related user data involved in the embodiments of this application are all obtained with the user's authorization. Furthermore, when the embodiments of this application are applied to specific products or technologies, the data used must be authorized or agreed to by the user, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0104] Figure 4This is a schematic diagram of a computer device structure provided in an embodiment of this application. The computer device 400 may include one or more central processing units (CPUs) 401 and a memory 405, in which one or more application programs or data are stored.
[0105] The memory 405 can be volatile or persistent storage. The program stored in the memory 405 can include one or more modules, each module including a series of instruction operations on the computer device. Furthermore, the central processing unit 401 can be configured to communicate with the memory 405 and execute the series of instruction operations stored in the memory 405 on the computer device 400.
[0106] Computer device 400 may also include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0107] The central processing unit 401 can perform the aforementioned... Figures 1 to 3 The specific operations performed by the computer device in the illustrated embodiment will not be described in detail here.
[0108] It should be noted that although the steps in the flowcharts of the various embodiments are drawn sequentially according to the arrows, unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the various embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0109] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0110] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0111] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0112] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0113] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0114] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the image generation model training method and the audio-based image generation method described above.
Claims
1. A training method for an image generation model, characterized in that, include: Acquire training data, which includes first audio data, second audio data, image data, and age and gender labels corresponding to the first audio data and image data; An audio classifier is trained using the age and gender labels in the training data and the first audio data. An image classifier is trained using the age and gender labels in the training data and the image data. The second audio data is processed based on a pre-trained image generation model to obtain a predicted image corresponding to the second audio data, wherein the first audio data and the second audio data are not completely identical; wherein the training process of the pre-trained image generation model includes: inputting the first audio data into an initial image generation model to obtain a predicted image corresponding to the first audio data; using the distance between the predicted image corresponding to the first audio data and the image data in the training data as a second training loss; adjusting the parameters of the initial image generation model based on the second training loss to obtain the pre-trained image generation model; The second audio data is input into the audio classifier to obtain the predicted human voice feature label corresponding to the second audio data, and the predicted image is input into the image classifier to obtain the predicted human image feature label corresponding to the predicted image. Based on the predicted human voice feature labels and the predicted human image feature labels, the parameters of the pre-trained image generation model are adjusted to obtain the trained image generation model.
2. The method according to claim 1, characterized in that, The step of adjusting the parameters of the pre-trained image generation model based on the predicted human voice feature labels and the predicted human image feature labels to obtain the trained image generation model includes: The distance between the predicted human voice feature labels and the predicted human image feature labels is used as the first training loss, and the parameters of the pre-trained image generation model are adjusted based on the first training loss to obtain the trained image generation model. or, The predicted human voice feature labels and the predicted human image feature labels are input into the matching model to obtain the degree of matching between the predicted human voice feature labels and the predicted human image feature labels; The matching degree between the predicted human voice feature labels and the predicted human image feature labels is used as the first training loss, and the parameters of the pre-trained image generation model are adjusted based on the first training loss to obtain the trained image generation model.
3. The method according to claim 1, characterized in that, The step of training an audio classifier using the age and gender labels from the training data and the first audio data includes: Input the first audio data into the initial audio classifier to obtain the predicted age and gender labels corresponding to the first audio data; The parameters of the initial audio classifier are adjusted based on the predicted age and gender labels corresponding to the first audio data and the distance between the age and gender labels in the training data to obtain the audio classifier; The step of training an image classifier using the age and gender labels in the training data and the image data includes: The image data is input into an initial image classifier to obtain the predicted age and gender labels corresponding to the image data; The parameters of the initial image classifier are adjusted based on the predicted age and gender labels corresponding to the image data and the distance between the age and gender labels in the training data to obtain the image classifier.
4. The method according to claim 1, characterized in that, The acquisition of training data includes: Acquire multiple existing audio data sets; A portion of the existing audio data is used as the first audio data, and image data corresponding to each first audio data input by the user is obtained, as well as age and gender labels corresponding to the first audio data and image data; The existing audio data other than the first audio data is identified as the second audio data.
5. The method according to claim 1, characterized in that, The predicted human voice feature labels include age and gender labels, and the predicted human image feature labels are consistent with the predicted human voice feature labels.
6. An audio-based image generation method, characterized in that, include: Acquire the audio data to be processed; The audio data to be processed is input into a trained image generation model to obtain an image corresponding to the audio data to be processed. The trained image generation model is trained based on the training method of the image generation model according to any one of claims 1 to 5.
7. The method according to claim 6, characterized in that, The method further includes: In response to a user-initiated instruction to generate a vocal image, at least one song is determined from the user's playlist and / or the songs sung by the user as the audio data to be processed.
8. A computer device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the training method of the image generation model according to any one of claims 1 to 5 or the audio-based image generation method according to any one of claims 6 to 7.
9. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed on the computer, cause the computer to perform the training method for the image generation model as described in any one of claims 1 to 5 or the audio-based image generation method as described in any one of claims 6 to 7.