Continuous emotion guided image-to-music generation method

By introducing cross-modal potential space and contrast learning technology of emotional sharing in the image-to-music generation method, the problems of blurring and subjectivity of image-to-music association are solved, and high-quality image-to-music generation is achieved.

CN119943010AInactive Publication Date: 2025-05-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510008486.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The task of image generation pure music faces ambiguity and subjectivity challenges, the correlation between images and music is unclear, and the lack of unified evaluation criteria, which makes model optimization difficult.

Method used

Using a continuous emotion-guided image-to-music generation method, by constructing images and music sample sets, using autoencoders and projectors to project features into the cross-modal potential space of emotion-sharing, contrast learning is performed to generate music.

Benefits of technology

It realizes the generation of pure music directly from natural images without relying on image titles or lyrics. The generated music has high emotional correlation with the image, high smoothness and quality, and strong practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943010A_ABST
    Figure CN119943010A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence. The invention provides a continuous emotion guided image-to-music generation method. According to the embodiment of the invention, an end-to-end framework is provided, pure music is directly generated from natural images, and dependence on image titles or lyrics is not needed. In consideration of fuzziness and subjectivity of a task, emotion is introduced as a medium for guiding a cross-modal conversion process. A plug-and-play model is provided, and images are converted into music works by utilizing comparative learning. It reduces the distance between images and music with similar emotions and the distance between images or music with similar emotions in the same modality, which is effective for processing continuous value tags. The music generated through the method is high in emotion association degree with the image, high in fluency quality and high in practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the field of artificial intelligence technology, and in particular to a method for generating music from images with continuous emotion guidance. Background Art

[0002] The task of generating pure music from images faces two major challenges: ambiguity and subjectivity. First, the ambiguity problem stems from the fact that there is no clear correspondence between the image content and the music melody, which makes the conversion from image to music more complicated than the generation of image to text. Therefore, some studies choose to use lyrics as an intermediary to connect the two modalities of image and music, rather than directly seeking a direct correspondence between them. Secondly, the subjective problem stems from the different interpretations of musical details by different individuals, which leads to the lack of a unified evaluation standard, making model optimization difficult. According to synesthesia theory, the connection between music and images is more reflected in the emotional level rather than low-level features.

[0003] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0004] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the invention

[0005] The purpose of the embodiments of the present disclosure is to provide a method for generating music from images with continuous emotion guidance, thereby overcoming one or more problems caused by the limitations and defects of the related art at least to a certain extent.

[0006] According to an embodiment of the present disclosure, a method for generating music from an image with continuous emotion guidance is provided, the method comprising: Constructing an image sample set and a music sample set; wherein the image sample set includes an image positive sample set and an image negative sample set, and the music sample set includes a music positive sample set and a music negative sample set; Inputting the image sample set and the music sample set into an image autoencoder and a music autoencoder respectively to extract image features and music features; Using an image projector and a music projector respectively to project the image features and the music features into a cross-modal latent space of emotion sharing, to obtain an image embedding vector and a music embedding vector; Performing contrastive learning in a cross-modal latent space according to the image embedding vector and the music embedding vector to obtain a first contrastive loss, a second contrastive loss, and a third contrastive loss; According to the first contrast loss, the second contrast loss, the third contrast loss and the reconstruction loss of the music back-projector, an overall loss function is obtained, and based on the overall loss function, the trained image projector, the music projector and the music back-projector are obtained; Inputting the image to be detected into the trained image autoencoder to extract features of the image to be detected; The trained image projector projects the image features to be generated into a cross-modal latent space of emotion sharing to generate a target image embedding vector, and generates a target music embedding vector based on the target image embedding vector; The target music embedding vector is projected into the music feature space using the trained music back-projector, and the target music embedding vector is decoded using a music decoder to obtain generated music.

[0007] Furthermore, the steps of constructing the image sample set and the music sample set include: Randomly select an image from the image dataset as an anchor point to obtain the image anchor ; Based on the image anchor , respectively select the image anchor from the music dataset The top 100 most similar sentiment scores music clips and the lowest music clips, and obtain the music positive sample set and the music negative sample set; Based on the image anchor , respectively select the anchor image from the image dataset The one with the highest sentiment similarity score images and the minimum images, and obtain the image positive sample set and the image negative sample set.

[0008] Further, the image projector comprises three layers of a first fully connected network connected in sequence, wherein the first fully connected network comprises a first batch of normalization, a first fully connected layer and a first ReLU activation function connected in sequence; The music projector includes two layers of a second fully connected network connected in sequence, wherein the second fully connected network includes a second batch normalization, a second fully connected layer, and a second ReLU activation function connected in sequence; The music back-projector includes two layers of a third fully-connected network connected in sequence, and the third fully-connected network includes a third batch normalization, a third fully-connected layer and a third ReLU activation function connected in sequence.

[0009] Further, the step of performing contrastive learning in a cross-modal latent space according to the image embedding vector and the music embedding vector to obtain a first contrast loss, a second contrast loss, and a third contrast loss includes: Obtaining a first score function between the image and the music according to the image embedding vector and the music embedding vector; According to the image positive sample set , the music positive sample set , the music negative sample set and the first score function to obtain the image positive sample set , music positive sample set and music negative sample set The first contrastive loss between , to perform contrastive learning between modalities in the cross-modal latent space; According to the embedding of any image x1 and the embedding of any image x2 Obtaining a second score function between the two images; According to the image positive sample set , the image negative sample set and the second score function to obtain the image positive sample set And the image negative sample set The second contrast loss between ; According to the embedding of any music y1 and embedding of any music y2 Get the third score function between the two pieces of music; According to the music positive sample set , the music negative sample set and the third score function to obtain the music positive sample set And the music negative sample set A third contrastive loss is proposed between the two models to perform contrastive learning within the modality in the cross-modal latent space.

[0010] Furthermore, the expression of the first score function is:

[0011] in, is the cosine similarity between image x and music y, x is the image, y is the music, is the temperature hyperparameter, is the image embedding vector, Embedding vectors for music; The expression of the first contrast loss is:

[0012] in, Representative Set The number of elements in a is the music positive sample set Music Negative Sample Set The union of; The expression of the second score function is:

[0013] in, is the embedding of image x1, is the embedding of image x2; The expression of the second contrast loss is:

[0014] in, , For the corresponding set or Zhongyu Different collections of images; The expression of the third score function is:

[0015] in, is the embedding of music y1, is the embedding of music y2; The expression of the third contrast loss is:

[0016] in, , Is the corresponding set or Zhongyu Different music collection.

[0017] Further, the step of obtaining an overall loss function according to the first contrast loss, the second contrast loss, the third contrast loss and the reconstruction loss of the music back-projector comprises: Obtaining a total loss in a cross-modal contrastive learning process according to the first contrastive loss, the second contrastive loss, and the third contrastive loss; The overall loss function is obtained according to the total loss and the reconstruction loss.

[0018] Furthermore, the total loss in the cross-modal contrastive learning process is:

[0019] The expression of the reconstruction loss is:

[0020] in, .

[0021] The expression of the overall loss function is:

[0022] in, is the reconstruction loss, is the total loss during cross-modal contrastive learning.

[0023] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects: In the embodiments of the present disclosure, through the above-mentioned continuous emotion-guided image-to-music generation method, on the one hand, an end-to-end framework is proposed to generate pure music directly from natural images without relying on image titles or lyrics. Considering the ambiguity and subjectivity of the task itself, emotions are introduced as a medium to guide the cross-modal conversion process. A plug-and-play model is proposed to convert images into musical works using contrastive learning. It reduces the distance between images and music with similar emotions, as well as the distance between images or music with similar emotions in the same modality, which is effective for processing continuous value labels. On the other hand, the music generated by this method has a high correlation with the image emotion, high fluency quality, and strong practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0025] Figure 1 A diagram showing the steps of a method for generating continuous emotion-guided images into music in an exemplary embodiment of the present disclosure; Figure 2 A schematic diagram showing the structure of an image projector in an exemplary embodiment of the present disclosure is shown; Figure 3 A schematic structural diagram of a music projector in an exemplary embodiment of the present disclosure is shown; Figure 4 A schematic diagram showing the structure of a music back-projector in an exemplary embodiment of the present disclosure is shown; Figure 5 A schematic diagram showing a method for generating continuous emotion-guided images into music in an exemplary embodiment of the present disclosure; Figure 6 An example diagram of the generation effect in the exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0027] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0028] This example embodiment provides a method for generating music from images with continuous emotion guidance. Figure 1 As shown in , the continuous emotion-guided image-to-music generation method may include: steps S101 to S108.

[0029] Step S101: constructing an image sample set and a music sample set; wherein the image sample set includes an image positive sample set and an image negative sample set, and the music sample set includes a music positive sample set and a music negative sample set; Step S102: inputting the image sample set and the music sample set into an image autoencoder and a music autoencoder respectively to extract image features and music features; Step S103: projecting the image feature and the music feature into a cross-modal latent space of emotion sharing using an image projector and a music projector respectively, to obtain an image embedding vector and a music embedding vector; Step S104: performing contrastive learning in a cross-modal latent space according to the image embedding vector and the music embedding vector to obtain a first contrastive loss, a second contrastive loss, and a third contrastive loss; Step S105: obtaining an overall loss function according to the first contrast loss, the second contrast loss, the third contrast loss and the reconstruction loss of the music back-projector, and obtaining the trained image projector, the music projector and the music back-projector based on the overall loss function; Step S106: inputting the image to be detected into the trained image autoencoder to extract features of the image to be detected; Step S107: the trained image projector projects the image features to be generated into the emotion-sharing cross-modal latent space to generate a target image embedding vector, and generates a target music embedding vector according to the target image embedding vector; Step S108: Use the trained music back-projector to project the target music embedding vector into the music feature space, and use a music decoder to decode the target music embedding vector to obtain generated music.

[0030] Through the above continuous emotion-guided image-to-music generation method, on the one hand, an end-to-end framework is proposed to generate pure music directly from natural images without relying on image captions or lyrics. Considering the ambiguity and subjectivity of the task itself, emotions are introduced as a medium to guide the cross-modal conversion process. A plug-and-play model is proposed to convert images into musical works using contrastive learning. It reduces the distance between images and music with similar emotions, as well as the distance between images or music with similar emotions within the same modality, which is effective for processing continuous value labels. On the other hand, the music generated by this method has a high correlation with the image emotion, high fluency quality, and strong practicality.

[0031] Next, we will refer to Figures 1 to 6 Each step of the above-mentioned continuous emotion-guided image-to-music generation method in this example implementation is described in more detail.

[0032] In step S101 and step S103, the image autoencoder (Img-Encoder and Img-Decoder) and the music autoencoder (Mus-Encoder and Mus-Decoder) are first trained to construct feature spaces in their respective modalities. Due to the modular plug-and-play design, various existing image and music autoencoders can be flexibly matched. Image features are obtained by encoding with Img-Encoder ,music Obtain music features through Mus-Encoder encoding These vectors retain the modality-specific features. Then, through two projectors (Img-Projector and Mus-Projector), they are projected into the emotion-shared cross-modal latent space to obtain the embedding vectors and The image projector contains three layers of fully connected networks, and the music projector contains two layers of fully connected networks. Each layer contains batch normalization and ReLU activation functions. For the specific architecture, refer to Figure 2 and Figure 3 .in, Figure 2 is a structural diagram of an image projector. Figure 3 This is a schematic diagram of the structure of the music projector. The formula for this process is defined as follows:

[0033] Subsequently, in order to ensure that the music embedding projected into the cross-modal latent space can be reconstructed back into the original feature space through the music de-projector for music generation, a corresponding reconstruction loss function is designed. The music de-projector adopts a two-layer fully connected neural network architecture, each layer is equipped with a batch normalization layer and a ReLU activation function. The specific architecture is referenced Figure 4 The mathematical expression of the reconstruction loss function is as follows:

[0034] In addition, in the emotion-sharing cross-modal latent space, contrastive learning is proposed to explore the emotion consistency between the two modalities and within a single modality, so that music and images with similar emotions are closer, while those with different emotions are farther away. In the inference stage, the image features are projected into the emotion-sharing cross-modal latent space to obtain the embedding vector , the feature is further projected back into the music feature space and Mus-Decoder is used to generate music. The proposed model is compatible with most image autoencoders and music autoencoders and has a plug-and-play nature.

[0035] In step S104, the selection of positive and negative samples for the contrastive learning algorithm in the cross-modal latent space. Multiple music clips may correspond to the emotion of a given image, so the music clips and images are grouped according to their VA values ​​in order to learn a more uniform emotional correlation between different music and images. The specific operation is to randomly select an image from the dataset as an anchor point. For a given image anchor , in order to construct a music positive sample set and music negative sample set , respectively select and The top 100 most similar sentiment scores music clips and the lowest The formula for this process is as follows:

[0036] in is a collection of all the music clips. is the image anchor, and Represents the top and bottom with the highest and lowest sentiment similarity scores, respectively. After Considering that there may be multiple images corresponding to the same emotion, we first randomly select an image from the dataset. As anchors. Then, by selecting the anchor images from the image dataset The one with the highest sentiment similarity score images and the minimum images, construct a positive sample set of images (denoted as ) and the image negative sample set (denoted as ). The formula for this process is as follows:

[0037] in is the set of all images, is the image anchor, and Represents the top and bottom with the highest and lowest sentiment similarity scores, respectively. After elements. After obtaining the positive sample sets of images and music ( , ) and negative sample set ( , ), we can leverage them for cross-modal and intra-modal contrastive learning. Cross-modal contrastive learning explores cross-modal sentiment consistency, while intra-modal contrastive learning enables us to learn more robust and discriminative latent vectors for each modality in the cross-modal latent space.

[0038] Inter-modality contrastive learning algorithm in cross-modal latent space. Given a set of positive image samples As an anchor point, the music positive sample set As positive samples, music negative sample set As a negative sample, define the image and music The score function between is:

[0039] in represents the cosine similarity, represents the temperature hyperparameter. is by using Img-Encoder from the image Extracted feature vectors is projected into the cross-modal latent space. Similarly, is by using Mus-Encoder to Extracted feature vectors Projected into the cross-modal latent space. Image positive sample set , music positive sample set and music negative sample set The contrast loss between them is calculated as:

[0040] in Representative Set The cross-modal contrast loss can shorten the distance between images and music that share similar emotions, while expanding the distance between those that are not similar in emotion, thereby better capturing the emotional relationship between images and music and generating emotionally consistent cross-modal content.

[0041] Intra-modality contrastive learning algorithm in cross-modal latent space. For image modality, given a set of positive image samples And image negative sample set , define the score function between two images as:

[0042] Image positive sample set And image negative sample set The contrast loss between them is:

[0043] in , Is the corresponding set or Zhongyu Different image collections. Specifically, if belong ,but Corresponds to Remove A collection of elements. Representative Set The number of elements in . The calculation of music modality is similar to that of image modality. Given a music positive sample set and music negative sample set , define the score function between two pieces of music as:

[0044] Music Positive Sample Set and music negative sample set The contrast loss between them is:

[0045] in , Is the corresponding set or Zhongyu Different music collections. Specifically, if belong ,but Corresponds to Remove A collection of elements. Representative Set The number of elements in .

[0046] In step S105, the overall loss function of the image-to-music generation algorithm. Given an image and music, the image and music are grouped into a positive sample set and a negative sample set respectively, and the total loss in the cross-modal contrastive learning process is:

[0047] Therefore, combined with the reconstruction loss, the total loss is defined as:

[0048] In step S106 to step S108, the image to be detected is input into the trained image autoencoder to extract the features of the image to be detected; the trained image projector projects the features of the image to be generated into the emotion-sharing cross-modal latent space to generate a target image embedding vector, and generates a target music embedding vector based on the target image embedding vector; the trained music back-projector is used to project the target music embedding vector into the music feature space, and the music decoder is used to decode the target music embedding vector to obtain the generated music.

[0049] like Figure 5 , which is a schematic diagram of the continuous emotion-guided image-to-music generation method.

[0050] In a specific embodiment, the image to be detected is input into a trained image autoencoder to extract the features of the image to be detected; the trained image projector projects the features of the image to be generated into a cross-modal latent space of emotional sharing to generate a target image embedding vector, and generates a target music embedding vector based on the target image embedding vector; the trained music back-projector is used to project the target music embedding vector into the music feature space, and the music decoder is used to decode the target music embedding vector to obtain the generated music. Figure 6 As shown, music that is consistent with the emotion of the image and has high fluency can be generated in the end.

[0051] Through the above continuous emotion-guided image-to-music generation method, on the one hand, an end-to-end framework is proposed to generate pure music directly from natural images without relying on image captions or lyrics. Considering the ambiguity and subjectivity of the task itself, emotions are introduced as a medium to guide the cross-modal conversion process. A plug-and-play model is proposed to transform images into musical works using contrastive learning. It reduces the distance between images and music with similar emotions, as well as the distance between images or music within the same modality, which is effective for processing continuous value labels. On the other hand, the music generated by this method has a high correlation with the image emotion, high fluency quality, and strong practicality.

[0052] It should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc. in the above description indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present disclosure and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the embodiments of the present disclosure.

[0053] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0054] In the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances.

[0055] In the embodiments of the present disclosure, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may include the first and second features being in direct contact, or may include the first and second features not being in direct contact but being in contact through another feature between them. Moreover, a first feature being "above", "above" and "above" a second feature includes the first feature being directly above and obliquely above the second feature, or simply indicates that the first feature is higher in level than the second feature. A first feature being "below", "below" and "below" a second feature includes the first feature being directly below and obliquely below the second feature, or simply indicates that the first feature is lower in level than the second feature.

[0056] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

[0057] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. A method for generating music from images with continuous emotion guidance, characterized in that: The method includes: Constructing an image sample set and a music sample set; wherein the image sample set includes an image positive sample set and an image negative sample set, and the music sample set includes a music positive sample set and a music negative sample set; Inputting the image sample set and the music sample set into an image autoencoder and a music autoencoder respectively to extract image features and music features; Using an image projector and a music projector respectively to project the image features and the music features into a cross-modal latent space of emotion sharing, to obtain an image embedding vector and a music embedding vector; Performing contrastive learning in a cross-modal latent space according to the image embedding vector and the music embedding vector to obtain a first contrastive loss, a second contrastive loss, and a third contrastive loss; According to the first contrast loss, the second contrast loss, the third contrast loss and the reconstruction loss of the music back-projector, an overall loss function is obtained, and based on the overall loss function, the trained image projector, the music projector and the music back-projector are obtained; Inputting the image to be detected into the trained image autoencoder to extract features of the image to be detected; The trained image projector projects the image features to be generated into a cross-modal latent space of emotion sharing to generate a target image embedding vector, and generates a target music embedding vector based on the target image embedding vector; The target music embedding vector is projected into the music feature space using the trained music back-projector, and the target music embedding vector is decoded using a music decoder to obtain generated music.

2. The method for generating continuous emotion-guided images into music according to claim 1, characterized in that: The steps of constructing the image sample set and the music sample set include: Randomly select an image from the image dataset as an anchor point to obtain the image anchor ; Based on the image anchor , respectively select the image anchor from the music dataset The top 100 most similar sentiment scores music clips and the lowest music clips, and obtain the music positive sample set and the music negative sample set; Based on the image anchor , respectively select the anchor image from the image dataset The one with the highest sentiment similarity score images and the minimum images, and obtain the image positive sample set and the image negative sample set.

3. The method for generating continuous emotion-guided images into music according to claim 1, characterized in that: The image projector includes three layers of a first fully connected network connected in sequence, wherein the first fully connected network includes a first batch of normalization, a first fully connected layer, and a first ReLU activation function connected in sequence; The music projector includes two layers of a second fully connected network connected in sequence, wherein the second fully connected network includes a second batch normalization, a second fully connected layer, and a second ReLU activation function connected in sequence; The music back-projector includes two layers of a third fully-connected network connected in sequence, and the third fully-connected network includes a third batch normalization, a third fully-connected layer and a third ReLU activation function connected in sequence.

4. The method for generating continuous emotion-guided images into music according to claim 1, characterized in that: The step of performing contrastive learning in a cross-modal latent space according to the image embedding vector and the music embedding vector to obtain a first contrast loss, a second contrast loss, and a third contrast loss includes: Obtaining a first score function between the image and the music according to the image embedding vector and the music embedding vector; According to the image positive sample set , the music positive sample set , the music negative sample set and the first score function to obtain the image positive sample set , music positive sample set and music negative sample set The first contrastive loss between , to perform contrastive learning between modalities in the cross-modal latent space; According to the embedding of any image x1 and the embedding of any image x2 Obtaining a second score function between the two images; According to the image positive sample set , the image negative sample set and the second score function to obtain the image positive sample set And the image negative sample set The second contrast loss between ; According to the embedding of any music y1 and embedding of any music y2 Get the third score function between the two pieces of music; According to the music positive sample set , the music negative sample set and the third score function to obtain the music positive sample set And the music negative sample set A third contrastive loss is proposed between the two models to perform contrastive learning within the modality in the cross-modal latent space.

5. The method for generating continuous emotion-guided images into music according to claim 4, characterized in that: The expression of the first scoring function is: in, is the cosine similarity between image x and music y, where x is the image and y is the music. is the temperature hyperparameter, is the image embedding vector, Embedding vectors for music; The expression of the first contrast loss is: in, Representative Set The number of elements in a is the music positive sample set Music Negative Sample Set The union of; The expression of the second score function is: in, is the embedding of image x1, is the embedding of image x2; The expression of the second contrast loss is: in, , For the corresponding set or Zhongyu Different collections of images; The expression of the third score function is: in, is the embedding of music y1, is the embedding of music y2; The expression of the third contrast loss is: in, , Is the corresponding set or Zhongyu Different music collection.

6. The method for generating continuous emotion-guided images into music according to claim 5, characterized in that: The step of obtaining an overall loss function according to the first contrast loss, the second contrast loss, the third contrast loss and the reconstruction loss of the music back-projector comprises: Obtaining a total loss in a cross-modal contrastive learning process according to the first contrastive loss, the second contrastive loss, and the third contrastive loss; The overall loss function is obtained according to the total loss and the reconstruction loss.

7. The method for generating continuous emotion-guided images into music according to claim 6, characterized in that: The total loss during cross-modal contrastive learning is: The expression of the reconstruction loss is: in, . The expression of the overall loss function is: in, is the reconstruction loss, is the total loss during cross-modal contrastive learning.