Audio generation method, device, computer equipment and storage medium

By extracting the image information of the target image and converting it into melody chord track note data, the audio file is directly synthesized, which solves the problem of low audio generation efficiency in traditional technology, and achieves efficient and high-quality audio generation.

CN114333744BActive Publication Date: 2025-05-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111327975.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-05-16
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

Traditional techniques are less efficient when training deep learning models and generating audio corresponding to target images, requiring a large number of training samples and a long training time.

Method used

By extracting the image information of each pixel point in the target image, determining the pitch value corresponding to the image information of each pixel point, converting it into melody track note data, and determining the chord track note data based on the melody track note data or the musical mode matching the target image, finally synthesize the melody track note data and chord track note data to generate an audio file.

Benefits of technology

This method can significantly improve the efficiency of audio generation, save time in collecting samples and model training, and the generated audio files have high quality and rich music effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333744B_ABST
    Figure CN114333744B_ABST
Patent Text Reader

Abstract

The present application relates to an audio generation method, device, computer equipment, storage medium and computer program product. The method can be applied to the application scenario of smart transportation, including: extracting the image information of each pixel in the target image; determining the pitch value corresponding to the image information of each pixel; converting the image information of each pixel into melody track note data based on the pitch value; determining the matching chord track note data based on the melody track note data or the music mode matching the target image; synthesizing the melody track note data with the chord track note data to obtain an audio file. The use of this method can improve the efficiency of generating audio files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an audio generation method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] With the development of computer technology, a large number of images need to be dubbed in order to synthesize the images and the dubbing of the images into a video. In traditional technology, a deep learning model is trained and the target image is processed by the trained deep learning model to obtain the audio corresponding to the target image. However, when training a deep learning model, a large number of training samples need to be collected, and the process of training a deep learning model takes a long time. Therefore, the efficiency of traditional technology in generating corresponding audio for a target image is low. Summary of the invention

[0003] Based on this, it is necessary to provide an audio generation method, apparatus, computer device, storage medium and computer program product that can improve generation efficiency in response to the above technical problems.

[0004] An audio generation method, the method comprising:

[0005] Extract image information of each pixel in the target image;

[0006] Determine a pitch value corresponding to the image information of each pixel;

[0007] Converting the image information of each pixel into melody track note data based on the pitch value;

[0008] Determining matching chord track note data based on the melody track note data or the music mode matching the target image;

[0009] The melody track note data and the chord track note data are synthesized to obtain an audio file.

[0010] An audio generation device, the device comprising:

[0011] An extraction module, used to extract image information of each pixel in the target image;

[0012] A determination module, used to determine the pitch value corresponding to the image information of each pixel point;

[0013] A conversion module, used for converting the image information of each pixel into melody track note data based on the pitch value;

[0014] The determination module is further configured to determine matching chord track note data based on the melody track note data or a music mode matching the target image;

[0015] The synthesis module is used to synthesize the melody track note data and the chord track note data to obtain an audio file.

[0016] In one embodiment, the note is a note in a chord; the device further comprises:

[0017] The determination module is further used to determine the music mode based on the image information of each pixel point;

[0018] A first selection module, used for selecting a chord template from a chord template library according to the music mode;

[0019] The determination module is further used to determine the chord series corresponding to the image information of each pixel point;

[0020] A first acquisition module, used for acquiring the notes in the chord corresponding to the chord series in the chord template;

[0021] The determination module is further used to determine the pitch value corresponding to the image information of each pixel point based on the notes in the chord.

[0022] In one embodiment, the apparatus further comprises:

[0023] The determination module is further used to determine the music mode based on the image information of each pixel point;

[0024] A second acquisition module is used to acquire a target note set corresponding to the music mode;

[0025] The determination module is further used to determine the pitch value corresponding to the image information of each pixel point based on each note in the target note set.

[0026] In one embodiment, the determining module is further used to:

[0027] In the chord template corresponding to the music mode, determining the first chord constituent tone corresponding to the image information of each pixel point, and generating chord track note data based on the first chord constituent tone; or,

[0028] Based on the melody note data of each section in the melody track note data, the matching chord note data is determined to obtain the chord track note data composed of the chord note data; or,

[0029] A second chord constituent tone that is fixedly matched with the music mode is obtained, and chord track note data is generated based on the second chord constituent tone.

[0030] In one embodiment, the determining module is further used to:

[0031] Selecting a chord template from a chord template library according to the music mode;

[0032] Determine the chord series corresponding to the image information of each pixel;

[0033] In the chord template, obtaining the chord constituent notes corresponding to the chord series;

[0034] The chord constituent tone corresponding to the chord series is determined as the first chord constituent tone corresponding to the image information of each pixel point.

[0035] In one embodiment, the determining module is further used to:

[0036] Determine the candidate chords corresponding to the melody note data of each section in the melody track note data;

[0037] Arrange and combine the candidate chords to obtain at least two candidate chord combinations;

[0038] In each of the alternative chord combinations, the chord degree corresponding to the current alternative chord is used as a reference level, and adjacent alternative chords are scored until the scores corresponding to all the alternative chords in each of the alternative chord combinations are obtained;

[0039] Determine a combined score of each of the candidate chord combinations based on the obtained scores;

[0040] Selecting a target chord combination from at least two candidate chord combinations based on the combination scores;

[0041] Chord note data matching the melody note data of each section is determined according to the target chord combination, and chord track note data composed of the chord note data is obtained.

[0042] In one embodiment, the determining module is further used to:

[0043] Determine the weight value corresponding to each note data in the melody note data of each section;

[0044] Weighting each of the note data according to the weight value to obtain a note score for each of the note data;

[0045] Based on the note values, determining the note and value of the chord corresponding to each of the chord progressions;

[0046] Sort the chords corresponding to the chord progressions according to the note sum values;

[0047] From the chords, the chords whose ranking reaches a preset ranking are selected as the candidate chords.

[0048] In one embodiment, the apparatus further comprises:

[0049] A normalization module, used for normalizing the melody track note data to obtain normalized melody track note data;

[0050] The determination module is further used to weight each note data in the normalized melody note data according to the weight value to obtain the note score of each note data.

[0051] In one embodiment, the image information is a brightness value; the extraction module is further used to:

[0052] Obtaining the chromaticity value of each pixel in the target image;

[0053] Determine the brightness value of each pixel in the target image based on the chromaticity value;

[0054] Determining the pitch value corresponding to the image information of each pixel point comprises:

[0055] Determine the pitch value corresponding to the brightness value of each pixel point.

[0056] In one embodiment, the melody track note data includes at least two sections of melody note data, and each section of the melody note data includes at least two note data; the device further includes:

[0057] A merging module, used for merging the same note data that appear continuously in the melody note data of each section to obtain merged melody note data;

[0058] The determination module is further used to determine the note value of each note data in the merged melody note data of each section;

[0059] The conversion module is further used to convert the image information of each pixel into melody track note data based on the pitch value and the note value.

[0060] In one embodiment, the apparatus further comprises:

[0061] A third acquisition module, used to acquire media materials and generate a target video based on the media materials;

[0062] A second selection module is used to select a target video frame from the target video as the target image;

[0063] An adjustment module, used for adjusting the aspect ratio of the target image according to the composition duration of the target video to obtain the adjusted target image;

[0064] The extraction module is further used to extract image information of each pixel in the adjusted target image.

[0065] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0066] Extract image information of each pixel in the target image;

[0067] Determine a pitch value corresponding to the image information of each pixel;

[0068] Converting the image information of each pixel into melody track note data based on the pitch value;

[0069] Determining matching chord track note data based on the melody track note data or the music mode matching the target image;

[0070] The melody track note data and the chord track note data are synthesized to obtain an audio file.

[0071] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0072] Extract image information of each pixel in the target image;

[0073] Determine a pitch value corresponding to the image information of each pixel;

[0074] Converting the image information of each pixel into melody track note data based on the pitch value;

[0075] Determining matching chord track note data based on the melody track note data or the music mode matching the target image;

[0076] The melody track note data and the chord track note data are synthesized to obtain an audio file.

[0077] A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the following steps are implemented:

[0078] Extract image information of each pixel in the target image;

[0079] Determine a pitch value corresponding to the image information of each pixel;

[0080] Converting the image information of each pixel into melody track note data based on the pitch value;

[0081] Determining matching chord track note data based on the melody track note data or the music mode matching the target image;

[0082] The melody track note data and the chord track note data are synthesized to obtain an audio file.

[0083] The above-mentioned audio generation method, apparatus, computer equipment, storage medium and computer program product convert the image information of each pixel point into melody track note data based on the pitch value corresponding to the image information of each pixel point in the target image. The chord track note data is determined based on the music mode matching the melody track note data or the target image. The melody track note data and the chord track note data are synthesized into an audio file. The audio file is obtained by processing the image information of each pixel point in the target image. Compared with obtaining the audio file through a deep learning model trained by a large number of training samples, the time for collecting samples and training models can be saved, thereby improving the efficiency of generating audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 A diagram showing an application environment of an audio generation method in an embodiment;

[0085] Figure 2 is a schematic flow chart of an audio generation method in one embodiment;

[0086] Figure 3 A schematic diagram of a flow chart of a method for determining a pitch value in one embodiment;

[0087] Figure 4 A schematic diagram of a flow chart of a method for generating chord track note data in one embodiment;

[0088] Figure 5 A schematic flow chart of a method for generating chord track note data in another embodiment;

[0089] Figure 6 A schematic diagram of various chord levels corresponding to the key of C major in an embodiment;

[0090] Figure 7 A schematic diagram of the correspondence between chords of various levels and melody note data of various sections in one embodiment;

[0091] Figure 8 A schematic diagram of a flow chart of a method for obtaining alternative chords in one embodiment;

[0092] Fig. 9 A schematic diagram of a flow chart of a method for generating chord track note data in one embodiment;

[0093] Fig.10 A schematic diagram of a flow chart of a method for generating melody track note data in one embodiment;

[0094] Fig.11 A schematic diagram of a flow chart of a method for extracting image information in one embodiment;

[0095] Fig.12 A schematic diagram of a flow chart of a method for determining a pitch value in one embodiment;

[0096] Fig.13 A schematic diagram of a process for generating an audio file in one embodiment;

[0097] Fig.14 A schematic diagram of a process for generating an audio file in another embodiment;

[0098] Fig.15 A schematic diagram of a process for generating an audio file in another embodiment;

[0099] Fig.16 A schematic diagram of a microservice architecture system in one embodiment;

[0100] Fig.17 is a structural block diagram of an audio generating device in one embodiment;

[0101] Fig.18 is a structural block diagram of an audio generating device in another embodiment;

[0102] Fig.19 is an internal structure diagram of a computer device in one embodiment;

[0103] Fig. 20 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION

[0104] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0105] The audio generation method provided in this application can be applied to Figure 1 In the application environment shown. In the application environment, a computer device 102 is included. The computer device 102 extracts image information of each pixel in the target image; determines a pitch value corresponding to the image information of each pixel; converts the image information of each pixel into melody track note data based on the pitch value; determines matching chord track note data based on the melody track note data or a music mode matching the target image; synthesizes the melody track note data and the chord track note data to obtain an audio file.

[0106] The computer device 102 may be a terminal or a server. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc. In addition, it may also be an intelligent voice interaction device, a smart home appliance, and a vehicle-mounted terminal, etc., but is not limited thereto.

[0107] The server can be an independent physical server or a server cluster composed of multiple service nodes in the blockchain system. Each service node forms a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol running on the Transmission Control Protocol (TCP).

[0108] In addition, the server can also be a server cluster composed of multiple physical servers, which can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.

[0109] In one embodiment, Figure 2 As shown, a method for generating an audio signal is provided. Figure 1 The computer device in the example is used to illustrate, including the following steps:

[0110] S202: extracting image information of each pixel in the target image.

[0111] The target image can be any color image or grayscale image, including a video frame extracted from a video, a picture captured from a web page, etc. The target image can be a compressed or uncompressed image.

[0112] Image information is information used to represent the characteristics of the target image, including image brightness, transparency (i.e., alpha channel information), chrominance, and digital fingerprints. For a compressed and coded target image, the computer device first decompresses the target image, and then extracts the image information of each pixel from the decompressed data corresponding to each pixel; for an uncompressed and coded target image, the computer device directly extracts the image information of each pixel from the data corresponding to each pixel in the target image.

[0113] S204, determining the pitch value corresponding to the image information of each pixel.

[0114] Among them, the pitch value is used to indicate the height of the sound, and can be represented by English letters, numbers, special symbols, etc. For example, the letters C, D, E, F, G, A, and B are used to represent the sounds of different pitches of do, re, mi, fa, sol, la, and si, respectively, and the pitch increases gradually from do to si; or the numbers 1, 2, 3, 4, 5, 6, and 7 can be used to represent the sounds of do, re, mi, fa, sol, la, and si, and the pitch increases gradually from 1 to 7.

[0115] The computer device can set the corresponding relationship between the image information and the pitch value, and determine the pitch value corresponding to the image information of each pixel according to the corresponding relationship. The corresponding relationship between the image information and the pitch value can be a linear or nonlinear relationship, for example, the pitch value is proportional to or inversely proportional to the image information, or the pitch value y=tan(image information), etc. For example, the image information is image transparency, and the image transparency value range is set to 0-100, 0 means that the image is completely opaque, and 100 means that the image is completely transparent. The pitch value is set to C or E. When the image transparency value is 0-50, the corresponding pitch value is determined to be C, and when the image transparency value is 51-100, the corresponding pitch value is determined to be E.

[0116] S206: Convert the image information of each pixel into melody track note data based on the pitch value.

[0117] Among them, the melody track note data is data used to record the melody generated based on the image information of each pixel point, which can be data in the MIDI (Musical Instrument Digital Interface) file format or data in other formats. A MIDI file is a binary data file consisting of a file header and a data description. In a MIDI format file, the pitch value is represented by a number, for example, the number 60 represents the pitch value C4. For example, the brightness values ​​of the eight pixels in the target image are 30, 74, 110, 120, 50, 180, 200, and 240 respectively, and the pitch values ​​determined based on the brightness values ​​of each pixel point are C4, E4, F4, F4, D4, A4, B4, and B4 respectively. If 60, 62, 64, 65, 69, and 71 are used in the MIDI file to represent the pitch values ​​C4, D4, E4, F4, A4, and B4, respectively, then the obtained melody track note data are 60, 64, 65, 65, 62, 69, 71, and 71.

[0118] In one embodiment, the melody track note data includes at least two sections of melody note data, and each section of melody note data includes at least two note data. Among them, the melody note data is the note data corresponding to a section of melody. The melody track note data can also include the sound value corresponding to each note data. The sound value is the note value, also known as the note value or sound value, which is used to indicate the relative duration between each note. The sound value of a perfect note is equal to the sound value of two half notes, or the sound value of four quarter notes, or the sound value of eight eighth notes, or the sound value of sixteen sixteenth notes, that is, if the duration of a perfect note is 1s, the duration of a half note is 1 / 2 second, the duration of a quarter note is 1 / 4 second, the duration of an eighth note is 1 / 8 second, and the duration of a sixteenth note is 1 / 16 second.

[0119] In one embodiment, the melody track note data may also include a force value corresponding to each note data. The computer device may determine the force value corresponding to each note data according to the image information of each pixel point in the target image.

[0120] The velocity value is used to indicate the strength of the sound, including very weak, weak, medium weak, medium strong, strong, very strong, sudden strong, etc., which can be represented by the characters pp, p, mp, mf, f, ff, sf respectively. The computer device can set a corresponding relationship between the image information and the velocity value, and based on the corresponding relationship, determine the velocity value corresponding to each note data according to the image information of each pixel.

[0121] The computer device determines the velocity value corresponding to each note data based on the image information of each pixel in the target image, so that the melody track note data includes the velocity value corresponding to each note data, thereby improving the sound effect of the audio synthesized based on the melody track note data and improving the efficiency of generating audio files.

[0122] S208: Determine matching chord track note data based on the melody track note data or the music mode matching the target image.

[0123] Among them, the musical mode is used to represent the organizational structure of notes. In the major and minor mode system, the musical mode can include major and minor keys. Major keys include natural major keys, harmonic major keys and melodic major keys, such as C major, D major or F major, etc.; minor keys include natural minor keys, harmonic minor keys and melodic minor keys, such as A minor.

[0124] Among them, a chord is a group of sounds with a certain interval relationship. Specifically, a chord is obtained by superimposing three or more notes according to a third or non-third interval relationship. For example, a chord can be composed of three notes C, E, and G, or it can also be composed of three notes D, F, and A. A chord composed of the first tone of the modal scale as the root note is a first-level chord, a chord composed of the second tone of the modal scale as the root note is a second-level chord, and so on. The chord track note data is data that records chords composed of notes, and can be data in MIDI file format.

[0125] The computer device may determine the matching chord note data based on the melody note data of each section in the melody track note data. For example, for the C major melody track note data including four sections of melody note data, the computer device determines that the chord matched by the first section of melody note data is a C chord, the chord matched by the second section of melody note data is an Am chord, the chord matched by the third section of melody note data is an F chord, and the chord matched by the fourth section of melody note data is a G chord. Then, the chord track note data is generated based on the C chord, the Am chord, the F chord, and the G chord.

[0126] In one embodiment, the melody track note data includes at least two sections of melody note data; the chord track note data includes chord note data corresponding to the melody note data of each section. The computer device determines the matching chord track note data based on the melody track note data in a fixed chord matching method and a predicted chord matching method. For the fixed chord matching method, the computer device matches the melody note data of each section with the pre-set chord note data, and obtains the chord track note data based on the chord note data of each section; for the predicted chord matching method, the computer device predicts the melody note data of each section to obtain the chord note data matching the melody note data.

[0127] In one embodiment, the computer device determines the matching chord track note data based on the music mode that matches the target image. For example, if the music mode that matches the target image is F sharp major, the computer device determines that the chords that match the four-section melody note data are B5 chord, Db5 chord, Eb5 chord, and Eb5 chord based on F sharp major, and then generates the corresponding chord track note data based on B5 chord, Db5 chord, Eb5 chord, and Eb5 chord.

[0128] S210, synthesizing the melody track note data and the chord track note data to obtain an audio file.

[0129] The audio file may be in various audio formats, including MP3 (Moving Picture Experts Group Audio Layer III), wav (Waveform Audio File Format), RealAudio and other formats.

[0130] In one embodiment, S210 specifically includes: the computer device merges the melody track note data with the chord track note data to obtain a MIDI format file; synthesizes the MIDI format file according to the sound source through the loaded synthesizer to obtain an audio file. The synthesizer may be an application that synthesizes the MIDI format file with the sound source, for example, the FluidSynth application. The sound source is the source of the sound, and may be a musical instrument that produces the sound, for example, the sound source may be a guitar, a piano, a saxophone, etc. For example, the computer device may synthesize the MIDI format file according to the sound effect of the guitar to obtain an audio file with the sound effect of the guitar.

[0131] In the above embodiment, the image information of each pixel in the target image is converted into melody track note data based on the pitch value corresponding to the image information of each pixel in the target image. The chord track note data is determined based on the music mode that matches the melody track note data or the target image. The melody track note data and the chord track note data are synthesized into an audio file. By processing the image information of each pixel in the target image to obtain an audio file, compared with obtaining an audio file through a deep learning model trained by a large number of training samples, the time for collecting samples and model training can be saved, thereby improving the efficiency of generating audio.

[0132] In one embodiment, the notes are notes within a chord, such as Figure 3 As shown, before S204, S302-S308 are also included, and S204 specifically includes S310.

[0133] S302: Determine the music mode based on the image information of each pixel.

[0134] The computer device can determine the style of the target image based on the image information of each pixel in the target image, and then determine the matching music mode based on the style of the target image. For example, if the computer device determines that the color of the target image is bright and rich based on the chromaticity value of each pixel in the target image, it can determine that the style of the target image is cheerful based on the chromaticity value, thereby determining that the matching music mode is a minor key. For another example, if the computer device determines that the target image is a Chinese-style image based on the image information of each pixel in the target image, it determines that the matching music mode is an ancient style mode.

[0135] In one embodiment, the computer device may process the image information of each pixel through a deep learning model to obtain the music mode. The deep learning model may be a convolutional neural network model, a residual convolutional neural network model, a recurrent neural network model, etc.

[0136] S304: Select a chord template from a chord template library according to the music mode.

[0137] The chord template is used to represent a fixed chord combination. Each degree of chord in the chord template includes the corresponding chord constituent notes and available notes within the chord. For example, chord template 1 is a combination of first-degree chords, sixth-degree chords, fourth-degree chords, and fifth-degree chords, wherein the constituent notes of the first-degree chord are C3, E3, G2, C2, and the available notes within the chord are C3, E3, G3, C4, E4, G4, C5, E5, G5, C6; the constituent notes of the sixth-degree chord are E3, C3, A2, A1, and the available notes within the chord are C3, E3, A3, C4, E4, A4, C5, E5, A5, C6; the notes of the fourth-degree chord are C3, A2, F2, F1, and the available notes within the chord are C3, F3, A3, C4, F4, A4, C5, F5, A5, C6; the notes of the fifth-degree chord are B2, G2, D2, G1, and the available notes within the chord are D3, G3, B3, D4, G4, B4, D5, G5, B5, D6.

[0138] The chord template library stores music modes and corresponding chord templates, and the computer device can select the corresponding chord template from the chord template library according to the music mode. For example, when the music mode is C major, the computer device selects chord template 1 from the chord template library, and when the music mode is A minor, the computer device selects chord template 2 from the chord template library.

[0139] S306: Determine the chord series corresponding to the image information of each pixel.

[0140] The chord series is used to represent the series of each chord in the chord template. For example, the chord template 1 includes four chords, and the series of the four chords are first, sixth, fourth and fifth, respectively.

[0141] In one embodiment, S306 specifically includes: determining the image row or image column to which each pixel belongs; determining the chord series corresponding to the image information of each pixel according to the arrangement order of the image row or the arrangement order of the image column. For example, the chords in the chord template 1 are arranged in the order of first-degree chord, sixth-degree chord, fourth-degree chord, and fifth-degree chord. When the pixel is the first row of pixels in the target image, the corresponding chord series is determined to be first degree; when the pixel is the second row of pixels in the target image, the corresponding chord series is determined to be sixth degree; when the pixel is the third row of pixels in the target image, the corresponding chord series is determined to be fourth degree; when the pixel is the fourth row of pixels in the target image, the corresponding chord series is determined to be fifth degree.

[0142] S308, obtaining the notes in the chord corresponding to the chord series in the chord template.

[0143] In the chord template, each chord of the series corresponds to a fixed note within the chord. In one embodiment, the computer device stores the notes within the chord corresponding to each chord series in the chord template in a data table, so that the notes within the chord corresponding to each chord series can be obtained by a table lookup method.

[0144] S310, determining a pitch value corresponding to the image information of each pixel based on the notes in the chord.

[0145] The computer device determines the pitch value corresponding to the image information of each pixel according to the correspondence between the image information and the notes in each chord. For example, if the image information is a brightness value, the brightness value ranges from 0 to 255, and the notes in the chord are C3, E3, G3, C4, and E4, the note in the chord corresponding to the brightness value of 120 is G3, and the pitch value corresponding to the pixel is determined to be G3.

[0146] In the above embodiment, the music mode is determined based on the image information of each pixel point, a chord template is selected from the chord template library according to the music mode and the chord series is determined, and in the chord template, the notes in the chord corresponding to the chord series are obtained and the pitch value corresponding to the image information of each pixel point is determined based on the notes in the chord. Therefore, the image information of each pixel point can be converted into melody track note data according to the pitch value, and an audio file is obtained based on the melody track note data, thereby improving the efficiency of generating audio files.

[0147] In one embodiment, before S204, it also includes: determining the music mode based on the image information of each pixel point; obtaining a target note set corresponding to the music mode; S204 specifically includes: determining the pitch value corresponding to the image information of each pixel point based on each note in the target note set.

[0148] The target note set is a note set including multiple notes. Notes are symbols used to represent sounds, and can be letters, numbers, a combination of letters and numbers, or special symbols. For example, the notes can be C, B, 2, 3, C3, D4, or Etc. For example, the target note set corresponding to C major may be “C3, D3, E3, G3, A3, C4, D4, E4, G4, A4, C5, D5, E5, G5, A5, C6”.

[0149] The computer device determines the note corresponding to the image information of each pixel according to the correspondence between the image information and the notes in the target note set, and uses the pitch value represented by each note in the target note set as the pitch value corresponding to the image information of the pixel. For example, the target note set is "C3, D3, E3, G3, A3, C4, D4, E4, G4, A4, C5, D5, E5, G5, A5, C6", and the image information is a brightness value of 78. It is determined that the note corresponding to the pixel in the target note set is D3, and the pitch value corresponding to the image information of the pixel is D3.

[0150] In the above embodiment, the computer device determines the music mode based on the image information of each pixel, obtains the target note set corresponding to the music mode, and determines the pitch value corresponding to the image information of each pixel based on each note in the target note set. Therefore, the image information of each pixel can be converted into melody track note data according to the pitch value, and an audio file is obtained based on the melody track note data, thereby improving the efficiency of generating audio.

[0151] In one embodiment, Figure 4 As shown, S208 specifically includes S402 or S404 or S406.

[0152] S402: Determine, in a chord template corresponding to the music mode, a first chord constituent tone corresponding to the image information of each pixel, and generate chord track note data based on the first chord constituent tone.

[0153] The first chord constituent tones are constituent tones of chords of all levels in the chord template. For example, in the chord template corresponding to C major, the first chord is C chord, and the constituent tones of C chord are C3, E3, G2, C2, then the first chord constituent tones include at least C3, E3, G2, C2.

[0154] In one embodiment, the computer device converts the constituent tones of the first chord into numbers according to the MIDI file format requirements, and generates chord track note data in the MIDI file format. For example, the computer device converts the constituent tones of the C chord C3, E3, G2, and C2 into numbers, and obtains chord track note data in the MIDI file format of 48 52 43 36.

[0155] In one embodiment, S402 specifically includes: selecting a chord template from a chord template library according to a music mode; determining a chord series corresponding to the image information of each pixel; obtaining chord constituent tones corresponding to the chord series in the chord template; and determining the chord constituent tones corresponding to the chord series as first chord constituent tones corresponding to the image information of each pixel.

[0156] The computer device can select a corresponding chord template from the chord template library according to the music mode, and the selected chord template includes multiple chords of different levels. For example, the computer device selects a chord template corresponding to C major in the chord template library, and the levels of the four chords in the chord template are one, six, four and five respectively. The computer device can determine the corresponding chord level according to the image row or image column to which each pixel belongs. For example, for the chord template corresponding to C major, if the pixel is the first row of pixels in the target image, it is determined that the chord level corresponding to the image information of the pixel is one; if the pixel is the second row of pixels in the target image, it is determined that the chord level corresponding to the image information of the pixel is six.

[0157] The computer device selects a chord template from a chord template library according to the music mode, and obtains the chord constituent tones corresponding to the chord series in the chord template, so that chord track note data can be generated according to the chord constituent tones, and the chord track note data and the melody track note data can be synthesized into an audio file, so that the sound effect of the obtained audio file is richer and the efficiency of generating audio files is improved.

[0158] S404, based on the melody note data of each section in the melody track note data, determine the matching chord note data, and obtain the chord track note data composed of the chord note data.

[0159] The melody track note data may include at least two sections of melody note data. For each section of melody note data, matching chord note data may be determined so that the chord note data serves as an accompaniment to the melody note data. The computer device may make predictions based on the melody note data of each section to obtain chord note data that matches the melody note data of the section, thereby forming the chord track note data from each chord note data. The computer device may also input the melody note data of each section into a machine learning model, and obtain chord note data that matches the melody note data through prediction by the machine learning model.

[0160] S406, obtaining a second chord constituent tone that is fixedly matched with the music mode, and generating chord track note data based on the second chord constituent tone.

[0161] The second chord constituent tones are the constituent tones of the chords that are fixedly matched with the music mode. For example, for F sharp major, the chords that are fixedly matched with the melody note data from the first section to the fourth section are B5 chord, Db5 chord, Eb5 chord and Eb5 chord respectively, the constituent tones of B5 chord are B2, F#3, B3, the constituent tones of Db5 chord are C#3, G#3, C#4, and the constituent tones of Eb5 chord are D#3, A#3, D#4.

[0162] In one embodiment, the computer device stores the music mode and the corresponding second chord constituent tones in a data table, and S406 specifically includes: the computer device searches the data table for the second chord constituent tones that are fixedly matched with the music mode, and generates chord track note data based on the second chord constituent tones.

[0163] In the above embodiment, the computer device generates chord track note data based on the first chord constituent tone; or determines matching chord note data based on the melody note data of each section in the melody track note data to obtain chord track note data composed of each chord note data; or generates chord track note data based on the second chord constituent tone. Thus, the synthesized audio file can have richer musical effects without manual intervention, and the efficiency of generating audio files is improved.

[0164] In one embodiment, Figure 5 As shown, S404 specifically includes the following steps:

[0165] S502, determining candidate chords corresponding to the melody note data of each section in the melody track note data.

[0166] The alternative chords may be all or part of the chords at all levels corresponding to the music mode, and the alternative chords corresponding to the melody note data of each section may be the same or different. Figure 6 As shown, the chords corresponding to C major are C chord, Dm chord, Em chord, F chord, G chord, and Am chord, and the alternative chords corresponding to the melody note data of each section can be one or more of C chord, Dm chord, Em chord, F chord, G chord, and Am chord. For example, the alternative chords corresponding to the melody note data of the first section are C chord, Dm chord, and G chord, and the alternative chords corresponding to the melody note data of the second section are F chord, Am chord, C chord, and G chord.

[0167] S504: Arrange and combine the candidate chords to obtain at least two candidate chord combinations.

[0168] The candidate chord combination is a chord combination composed of chords selected from the candidate chords corresponding to the melody note data of each section. The computer device arranges and combines the candidate chords corresponding to the melody note data of each section to obtain at least two candidate chord combinations.

[0169] In one embodiment, S504 specifically includes: the computer device arranges the alternative chords corresponding to the melody note data of each section; selects a chord from each group of arranged alternative chords in turn; and composes the selected chords into an alternative chord combination. For example, the melody track note data includes three sections of melody note data, the alternative chords corresponding to the first section of the melody note data are 1st and 4th degree chords, the alternative chords corresponding to the second section of the melody note data are 2nd and 5th degree chords, and the alternative chords corresponding to the third section of the melody note data are 3rd and 4th degree chords, then the obtained alternative chord combinations are "123", "124", "153", "154", "423", "424", "453", "454".

[0170] S506: In each candidate chord combination, the chord degree corresponding to the current candidate chord is used as a reference level to score the adjacent candidate chords until the scores corresponding to all the candidate chords in each candidate chord combination are obtained.

[0171] The computer device scores the next section of alternative chords adjacent to the current alternative chord using the chord degree corresponding to the current alternative chord as a reference level. For example, the alternative chord combination is a chord combination consisting of a first-level chord, a fifth-level chord and a third-level chord. If the current chord is a first-level chord, the fifth-level chord is scored using the first-level chord as a reference level, and then the third-level chord is scored using the fifth-level chord as a reference level.

[0172] In one embodiment, the computer device scores the candidate chords according to a scoring table, which includes the chord degree corresponding to the current candidate chord and the scores corresponding to each adjacent candidate chord of the current candidate chord. For example, the scoring table is shown in Table 1. If the candidate chord combination is a chord combination consisting of a first-degree chord, a fifth-degree chord and a third-degree chord, the first-degree chord is used as the current candidate chord, and the fifth-degree chord is scored according to the scoring table shown in Table 1, then the score is 10, and the fifth-degree chord is used as the current candidate chord to score the third-degree chord, then the score is 8.

[0173] Table 1

[0174]

[0175] S508: Determine the combined score of each candidate chord combination based on the obtained score.

[0176] The computer device determines the combined scores of each candidate chord combination based on the scores obtained by scoring the candidate chords. In one embodiment, the computer device determines the sum of the obtained scores as the combined scores of each candidate chord combination; or the obtained scores may be weighted and summed, and the weighted sum is used as the combined scores of each candidate chord combination. For example, if the candidate chord combination is a chord combination consisting of a first-degree chord, a fifth-degree chord, and a third-degree chord, the first-degree chord is used as the current candidate chord, the score obtained by scoring the fifth-degree chord is 8, and the score obtained by scoring the third-degree chord with the fifth-degree chord as the current candidate chord is 10, then the combined score of the candidate chord combination is 18.

[0177] S510: Select a target chord combination from at least two candidate chord combinations based on the combination scores.

[0178] The target chord combination may be one or more chord combinations selected from at least two candidate chord combinations according to the combination score. For example, the target chord combination may be a chord combination whose combination score reaches a preset score, or may be a chord combination with the largest combination score.

[0179] In one embodiment, the computer device may sort at least two candidate chord combinations according to the combination scores, and select a chord combination ranked within a preset ranking from the sorted candidate chord combinations as the target chord combination.

[0180] S512, determining the chord note data matching the melody note data of each section according to the target chord combination, and obtaining the chord track note data composed of the chord note data.

[0181] The computer device matches each chord in the target chord combination with the melody note data of each section. In one embodiment, S512 specifically includes: the computer device matches the chords in the target chord combination with the melody note data of each section according to the arrangement order of the chords of each level in the target chord combination, and obtains the chord note data matching the melody note data of each section according to the chords matching the melody note data of each section. Figure 7 As shown, the target chord combination includes the first-degree chord, the fifth-degree chord, the third-degree chord and the sixth-degree chord arranged in order, and the computer device matches the first-degree chord with the first section melody note data; matches the fifth-degree chord with the second section melody note data; matches the third-degree chord with the third section melody note data; matches the sixth-degree chord with the fourth section melody note data, according to the chord track note data composed of the chord note data corresponding to the first-degree chord, the fifth-degree chord, the third-degree chord and the sixth-degree chord.

[0182] In the above embodiment, the candidate chords corresponding to the melody note data of each section in the melody track note data are determined and the candidate chords are arranged and combined to obtain at least two candidate chord combinations. Among the at least two candidate chord combinations, a target chord combination is selected based on the combination scores corresponding to the candidate chord combinations, and the chord note data matching the melody note data of each section is determined according to the target chord combination, so as to obtain the chord track note data composed of the chord note data. Therefore, the matching chord note data can be determined according to the melody track note data to obtain the chord track note data, so that the chord track note data can be obtained without manual intervention, thereby improving the efficiency of generating audio files.

[0183] In one embodiment, Figure 8 As shown, S502 specifically includes the following steps:

[0184] S802, determining the weight values ​​corresponding to the respective note data in the melody note data of each section.

[0185] The melody note data is the note data corresponding to a section of the melody. Each section of the melody includes multiple note data. For example, 1 1 5 5|6 6 5-|4 4 3 3|2 2 1-| is four sections of melody note data, "1 1 5 5" is the first section of the melody note data, and "1", "1", "5", "5" are the note data in the section of the melody note data.

[0186] In one embodiment, the computer device determines the weight values ​​corresponding to each note data according to the beat corresponding to the melody note data. The beat is used to represent the combination of strong beats or weak beats corresponding to each note data in the melody note data, including 1 / 4 beat, 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, etc. For example, for the 4 / 4 beat melody track note data, each section of the melody note data includes four quarter notes, and the note data of each quarter note corresponds to a strong beat, a weak beat, a secondary strong beat and a weak beat. The weight value corresponding to the note data of the strong beat can be greater than the weight value corresponding to the note data of the weak beat. For example, for the first measure 1 1 5 5 in 11 5 5|6 6 5-|4 4 3 3|22 1-|, the first "1" is a strong beat, and the corresponding weight value can be set to 0.4, the second "1" is a weak beat, and the corresponding weight value can be set to 0.2, the first "5" is the secondary strong beat, and the corresponding weight value can be set to 0.3, and the second "5" is a weak beat, and the corresponding weight value can be set to 0.1.

[0187] S804, weighting each note data according to the weight value to obtain the note score of each note data.

[0188] The computer device weights each note data in the melody track note data, sums the weighted scores, and obtains the note scores of each note data. For example, for CCGG|AA G-|FF EE|DD C-|, the weight value corresponding to the first "C" in the first measure is 0.4, the weight value corresponding to the second "C" in the first measure is 0.2, and the weight value corresponding to the "C" in the fourth measure is 0.3, then the note score of "C" is 100×0.4+100×0.2+100×0.3=90; the weight value corresponding to the first "G" in the first measure is 0.3, the weight value corresponding to the first "G" in the first measure is 0.1, and the weight value corresponding to the "G" in the second measure is 0.3, then the note score of "G" is 100×0.3+100×0.1+100×0.3=70.

[0189] S806: Determine the note and value of the chord corresponding to each chord series based on the note scores.

[0190] The chord corresponding to each chord series is composed of multiple chord constituent notes, each chord constituent note can be represented by corresponding note data, and the note sum value of the chord can be determined according to the note score corresponding to each note data. For example, the first-order chord corresponding to C major, that is, the C chord, is composed of chord constituent notes represented by three note data of C, E, and G. The note sum value of the C chord can be determined according to the note scores corresponding to C, E, and G respectively. For example, the note score corresponding to C is 10, the note score corresponding to E is 20, and the note score corresponding to G is 40, then the note sum value corresponding to the C chord is 70.

[0191] S808, sorting the chords corresponding to the chord degrees according to the note sum values; and selecting the chords whose sorted ranking reaches a preset ranking as candidate chords from the chords.

[0192] The computer device can sort the chords corresponding to each chord degree in the order of note sum from high to low or from low to high, and select the chords whose ranking reaches the preset ranking from the sorted chords as the alternative chords. For example, the sorted chords are 5th-degree chords, 4th-degree chords, 1st-degree chords, 6th-degree chords, 3rd-degree chords, and 2nd-degree chords, and the computer device can select the chords that reach the preset ranking as the alternative chords, for example, select the first 3 chords as the alternative chords, or select the chords ranked in the top 10% as the alternative chords, or select the chords whose note sum is greater than the preset value as the alternative chords.

[0193] In one embodiment, before S604, the process further includes: normalizing the melody track note data to obtain normalized melody track note data. S604 specifically includes: weighting each note data in the normalized melody note data according to the weight value to obtain the note score of each note data.

[0194] Among them, the normalization process is to normalize the note data of different pitches to the same pitch. For example, the results of normalizing the note data C1, C2, C3 and C4 are all C. After the computer device normalizes each note data in the melody track note data, it weights each note data after normalization according to the weight value corresponding to each note data to obtain the note score of each note data. For example, for the 4 / 4 beat melody track note data C3 A2 A2 E2|C4 A3 B2 D4|E3 D2 C2 F4|, the normalized melody track note data obtained by normalizing the melody track note data is CAAE|CABD|EDCF|, and for the melody note data of each section in the normalized melody track note data, the weight values ​​corresponding to the first to fourth beat note data are 0.4, 0.2, 0.3, and 0.1, respectively. The note score corresponding to C is 100×0.4+100×0.4+100×0.3=110, and the note score corresponding to A is 100×0.2+100×0.3+100×0.2=70.

[0195] In one embodiment, Fig. 9 As shown, S404 specifically includes the following steps:

[0196] S902, normalizing the melody note data of each section in the melody track note data.

[0197] S904, determining the weight values ​​corresponding to the respective note data in the note data of each section of the melody obtained after the normalization process.

[0198] S906, weighting each note data according to the weight value to obtain the note score of each note data.

[0199] S908: Determine the note and value of the chord corresponding to each chord series based on the note scores.

[0200] S910, sorting the chords corresponding to each chord degree according to the note sum value.

[0201] S912: Select, from among the chords, a chord whose ranking reaches a preset ranking as a candidate chord.

[0202] S914: Arrange and combine the candidate chords to obtain at least two candidate chord combinations.

[0203] S916, in each candidate chord combination, the chord degree corresponding to the current candidate chord is used as a reference level to score the adjacent candidate chords until the scores corresponding to all the candidate chords in each candidate chord combination are obtained.

[0204] S918: Determine a combined score of each candidate chord combination based on the obtained score.

[0205] S920: Select a target chord combination from at least two candidate chord combinations based on the combination scores.

[0206] S922, determining the chord note data that matches the melody note data of each section according to the target chord combination, and obtaining the chord track note data composed of the chord note data.

[0207] For details of S902 to S922, please refer to Figure 5-8 The specific implementation process in the embodiment.

[0208] In the above embodiment, the weight values ​​corresponding to the respective note data in the melody note data of each section are determined, and the note data are weighted according to the weight values ​​to obtain the note scores of the respective note data. Then, based on the note scores, the note sum values ​​of the chords corresponding to the chord series are determined, and the candidate chords are selected from the chords according to the note sum values. Thus, the candidate chords with a high degree of matching with the melody note data of each section can be obtained, and the candidate chord combinations can be obtained based on the candidate chords, and then the chord track note data can be generated based on the chord combinations selected from the candidate chord combinations, thereby improving the efficiency of generating audio files.

[0209] In one embodiment, Fig.10 As shown, the melody track note data includes at least two sections of melody note data, and each section of the melody note data includes at least two note data; before S206, it also includes S1002-S1004, and S206 specifically includes S1006.

[0210] S1002, in the melody note data of each section, merge the consecutive identical note data to obtain merged melody note data.

[0211] The merging process is to merge multiple consecutive identical note data into one or more note data. For example, if the melody note data is AACBEFFGDDDDCCG, the merged melody note data is ACBEFFGDCG.

[0212] In one embodiment, when the number of consecutive identical note data is greater than a preset value or the sum of the sound values ​​of consecutive identical note data is greater than a preset sound value, some of the identical note data are merged, and the remaining note data are converted to empty notes. When the number of empty notes obtained by the conversion reaches a preset value, the merging of consecutive identical note data is restarted. For example, when the number of consecutive identical note data is greater than 2, the first two consecutive identical note data can be merged, and the remaining note data can be converted to empty notes to obtain merged melody note data. For example, the melody note data is AACBEFFGDDDDDDCCG, and the merged melody note data is ACBEFFFGDDDDDDCCG. (blank note)GD (blank note)DCG.

[0213] S1004, determining the note value of each note data in the merged melody note data of each section.

[0214] The note value is a note duration, also known as a note value or a note value, and is used to indicate the relative duration between notes. The computer device can set the note value of the note data corresponding to each pixel, for example, setting the note data corresponding to each pixel to a 16th note, that is, the note value corresponding to the note data corresponding to each pixel is 1 / 16 of a full note.

[0215] In one embodiment, in each section of the melody note data, when the same note data that appear continuously are merged, the sound values ​​of the same note data that appear continuously are merged to obtain the sound value of each note data in the merged melody note data. For example, for two consecutive 16th notes, the two consecutive 16th notes are merged into an 8th note, and the sound value of the note data obtained after the merger is 1 / 8 of the whole note. For example, for four consecutive 16th notes, the four consecutive 16th notes are merged into a 4th note, and the sound value of the note data obtained after the merger is 1 / 4 of the whole note.

[0216] S1006: Based on the pitch value and the note value, the image information of each pixel is converted into melody track note data.

[0217] The computer device determines the pitch value and the note value corresponding to the image information of each pixel, and then obtains the melody track note data according to the pitch value and the note value. For example, the brightness value of the pixel is 120, the pitch value determined according to the brightness value is C4, and the corresponding note value is 1 / 16 of a full note. Assuming that 1 / 16 of a full note is 100ms, the melody track note data includes the note data 60 and the note value 100ms corresponding to C4.

[0218] In the above embodiment, in the melody note data of each section, the same note data that appear continuously are merged, and then the sound value of each note data in the merged melody note data of each section is determined, and based on the pitch value and the sound value, the image information of each pixel point is converted into melody track note data. Therefore, an audio file can be generated based on the melody track note data, which improves the efficiency of generating audio files.

[0219] In one embodiment, Fig.11 As shown, before S202, steps S1102-S1106 are also included, and S202 specifically includes S1108.

[0220] S1102: Acquire media materials, and generate a target video based on the media materials.

[0221] The media material is a multimedia material, including text material, picture material or video material, etc. The media material can be a material captured from a web page, a material obtained from a database, or a material uploaded by a client, etc. The target video can be a video in various formats such as DVD, MPEG-4, H.264, AVI, etc.

[0222] In one embodiment, S302 specifically includes: the computer device can crawl web page data from the web page through an application program, extract media materials from the web page data, and synthesize the extracted media materials into a target video. The application program can be, for example, a CROSS application program.

[0223] S1104, selecting a target video frame from the target video as a target image.

[0224] The target video frame may be one or more frames in the target video, may be the first video frame in the target video, or may be a video frame randomly selected from the target video. For compressed video, the target video frame may be an I frame (intra-coded frame), a B frame (inter-coded frame), or a P frame (forward prediction frame).

[0225] S1106, adjusting the aspect ratio of the target image according to the composition duration of the target video to obtain an adjusted target image.

[0226] The composition duration is the duration of dubbing in the target video, which may be the same as the duration of the target video or, when only part of the target video is dubbed, may be less than the duration of the target video.

[0227] In one embodiment, S1106 specifically includes: determining the number of sections of the melody note data corresponding to the composition duration of the target video; scaling the target image based on the number of sections of the melody note data, so as to adjust the aspect ratio of the target image through the scaling operation, so that the number of rows of the adjusted target image is the same as or greater than the number of sections of the melody note data. For example, the composition duration of the target video is 5 minutes, and the 5-minute audio includes 20 sections of melody note data. The computer device adjusts the number of rows of the target image to 20 rows or greater than 20 rows by scaling the target image.

[0228] S1108, extracting image information of each pixel in the adjusted target image.

[0229] After adjusting the aspect ratio of the target image, the computer device extracts the image information of each pixel in the adjusted target image, then determines the pitch value corresponding to each image information, and converts the image information of each pixel into melody track note data based on the pitch value.

[0230] In the above embodiment, the computer device generates a target video based on the acquired media material, and selects a target video frame from the target video as a target image. Then, the aspect ratio of the target image is adjusted according to the composition duration of the target video, and the image information of each pixel is extracted from the adjusted target image. Thus, the dubbing of the target video can be generated according to the image information of each pixel, which improves the efficiency of generating audio, and the audio file corresponding to the target video can be generated without human participation, thereby realizing the automatic generation of the audio file.

[0231] In one embodiment, Fig.12 As shown, the image information is a brightness value, S202 specifically includes S1202-S1204, and S204 specifically includes S1206.

[0232] S1202, obtaining the chromaticity value of each pixel in the target image.

[0233] Among them, brightness is used to indicate the brightness of a pixel. The higher the brightness value, the brighter the pixel. For digital images, the brightness of a pixel can be represented by a value of 0-255. Chroma is used to represent the color characteristics of a pixel. For the RGB color mode, the color characteristics of a pixel are represented by three color channels: R (Red), G (Green) and B (Blue), and the chroma values ​​are the values ​​of the three color channels: R, G and B; for the HSV color mode, the color characteristics of a pixel are represented by three parameters: H (Hue), S (Saturation) and V (Value), and the chroma values ​​are the values ​​of the three parameters: H, S and V. In one embodiment, the target image is an image in the RGB color mode.

[0234] S1204: Determine the brightness value of each pixel in the target image based on the chromaticity value.

[0235] Since there is a certain mapping relationship between the brightness value and the chromaticity value of a pixel, the computer device can determine the brightness value of the pixel based on the chromaticity value of a certain pixel according to the mapping relationship. For example, the computer device can perform a weighted summation of the chromaticity values ​​of each pixel to obtain the brightness value corresponding to the pixel. Assuming that Gray represents the brightness value of the pixel, and R, G, and B represent the chromaticity values ​​of the three color channels of the pixel, Gray = 0.30R + 0.59G + 0.11B.

[0236] S1206, determining the pitch value corresponding to the brightness value of each pixel.

[0237] The computer device can set the correspondence between the brightness value and the pitch value, and determine the pitch value corresponding to the brightness value of each pixel point according to the correspondence. For example, assuming that the brightness value is a value of 0-255, the pitch values ​​are C, D, E, F, G, A, and B. The correspondence between the brightness value and the pitch value is shown in Table 2. When the brightness value is 80, the computer device determines that the pitch value corresponding to the brightness value is E according to the correspondence shown in Table 2; when the brightness value is 150, the computer device determines that the pitch value corresponding to the brightness value is G according to the correspondence shown in Table 2.

[0238] Table 2

[0239] Brightness value 0-35 36-70 71-105 106-140 141-175 176-210 211-255 Pitch value C D E F G A B

[0240] In one embodiment, Fig.13 As shown, the computer device obtains media material and then generates a target video based on the media material. The computer device selects a target video frame from the target video as a target image and preprocesses the target image. After preprocessing the target image, image information is extracted from the preprocessed target image, and melody and chord composition are performed based on the image information. Melody track note data is generated through melody composition, and chord track note data is generated through chord composition. The melody track note data and the chord track note data are combined into a MIDI file, and then the MIDI file is synthesized into an audio file through a MIDI synthesizer, and the audio file and the target video are synthesized to obtain a target video with dubbing.

[0241] In one embodiment, the target image is an image of size 16×8, containing 16 columns and 8 rows of pixels. The computer device generates a corresponding section of melody note data based on each row of pixels, and determines the chord note data corresponding to the section of melody note data. Therefore, the computer device generates a total of 8 sections of melody note data and 8 sections of chord note data based on the target image.

[0242] like Fig.14 As shown, the computer device extracts the image information of each pixel in the target image, determines the music mode based on the image information of each pixel, and then selects a chord template from the chord template library based on the music mode. For example, the chord template is a chord combination consisting of a first chord, a sixth chord, a fourth chord, and a fifth chord.

[0243] The computer device determines the chords corresponding to the pixels in each row according to the arrangement order of the image rows to which the pixels belong and the arrangement order of the chords of each level in the chord template. The chords corresponding to the pixels in the first to eighth rows are the first chord, the sixth chord, the fourth chord, the fifth chord, the first chord, the sixth chord, the fourth chord, and the fifth chord, respectively. Based on the notes in the chords corresponding to the pixels in each row, the pitch values ​​corresponding to the image information of each pixel are determined. The note data corresponding to each pixel is determined based on the pitch values, and a section of melody note data is generated based on the note data corresponding to each row of pixels. In each section of the melody note data, the same note data that appears continuously is merged to obtain the merged melody note data. The sound value of each note data in the merged melody note data of each section is determined, and the image information of each pixel is converted into melody track note data based on the pitch value and the sound value.

[0244] The computer device obtains the chord constituent tones corresponding to the chords of the chord progression in the chord template according to the chord progression corresponding to each row of pixel points, and generates chord track note data according to the chord constituent tones. The melody track note data and the chord track note data are combined to obtain a MINI format file, and then the MINI file is synthesized with the sound source through a synthesizer to obtain an audio file.

[0245] In one embodiment, the target image is an image of size 16×16, containing 16 columns and 16 rows of pixels. The computer device generates a corresponding section of melody note data based on each row of pixels, and predicts the corresponding chord note data based on the section of melody note data, or fixes the corresponding chord note data to the section of melody note data.

[0246] like Fig.15 As shown, the computer device extracts the image information of each pixel in the target image, and determines the music mode based on the image information of each pixel. Then, it is selected whether to use the predicted chord mode or the fixed chord mode. The computer device can randomly select the predicted chord mode or the fixed chord mode, or the computer device can also select according to the current task amount. For example, if the number of target images that need to be processed currently exceeds the preset number threshold, the computer device can select the fixed chord mode because the calculation amount of the fixed chord mode is small.

[0247] The computer device obtains a target note set corresponding to the music mode (e.g., C3, D3, E3, G3, A3, C4, D4, E4, G4, A4, C5, D5, E5, G5, A5, C6), determines the pitch value corresponding to the image information of each pixel based on each note in the target note set, and determines the note data corresponding to each pixel based on the pitch value. For each row of pixels, if the pitch values ​​corresponding to adjacent pixels are the same, the note data corresponding to the adjacent pixels are merged to obtain merged note data. The note value of each note data in the merged melody note data of each section is determined, and the melody track note data is generated based on the pitch value and the note value.

[0248] For predicting chord patterns, the computer device normalizes each section of melody note data in the melody track note data. Determine the weight values ​​corresponding to each note data in each section of melody note data obtained after normalization. Weight each note data according to the weight value to obtain the note score of each note data. Based on the note score, determine the note sum value of the chord corresponding to each chord degree. Sort the chords corresponding to each chord degree according to the note sum value, select the chords with a preset ranking from each chord as alternative chords, and arrange and combine the alternative chords to obtain at least two alternative chord combinations. In each alternative chord combination, take the chord degree corresponding to the current alternative chord as the reference level, score the adjacent alternative chords, until the scores corresponding to all alternative chords in each alternative chord combination are obtained. Determine the combined score of each alternative chord combination based on the obtained score, and select the target chord combination based on the combined score in at least two alternative chord combinations. Chord note data matching the melody note data of each section is determined according to the target chord combination, and chord track note data composed of each chord note data is obtained.

[0249] For the fixed chord mode, the computer device obtains 4 chords that are fixedly matched with the music mode, and then loops the 4 chords to obtain 8 chords, and generates corresponding chord track note data for each chord.

[0250] The computer device combines the melody track note data with the chord track note data to obtain a file in MINI format, and then synthesizes the MINI file with a sound source through a synthesizer to obtain an audio file.

[0251] In one embodiment, the computer device is a Linux server, and the Linux server, Redis database, MySQL database, FluidSynth software package and gRPC protocol form a microservice architecture system. Fig.16As shown, the microservice architecture system includes an access layer, a service layer, a data layer, and an architecture layer. The access layer obtains user data and media data, verifies the user data and media data, and then provides the verified media data to the service layer. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0252] The service layer includes image service, audio service and video service. Image service includes image preprocessing and image analysis. Image preprocessing includes adjusting the image aspect ratio, color correction of the image, etc. Image analysis includes extracting image information from the image, and generating melody track note data and chord track note data based on the image information. Audio service includes combining melody track note data and chord track note data into MIDI files, and synthesizing MIDI files into audio files through FluidSynth software. The audio file and image are rendered through the video service to obtain a video file. The data layer includes image data cache, audio data cache, MIDI data cache, video data cache and video rendering data cache.

[0253] In one embodiment, the present application also provides an application scenario of smart transportation, which applies the above-mentioned audio generation method. Specifically, the application of the audio generation method in this application scenario is as follows: During driving, the vehicle-mounted terminal obtains the target image, extracts the image information of each pixel in the target image, and determines the pitch value corresponding to the image information of each pixel. Then, based on the pitch value, the image information of each pixel is converted into melody track note data. Based on the melody track note data or the music mode that matches the target image, the matching chord track note data is determined. The melody track note data and the chord track note data are synthesized to obtain an audio file. The vehicle-mounted terminal synthesizes the acquired target image and the audio file into a video and plays it.

[0254] It should be understood that although Figure 2-5 The steps in the flowcharts of 8-12 are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-5, at least part of the steps in 8-12 may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0255] In one embodiment, Fig.17 As shown, an audio generation device is provided, which can be a part of a computer device using a software module or a hardware module, or a combination of the two. The device specifically includes: an extraction module 1702, a determination module 1704, a conversion module 1706 and a synthesis module 1708, wherein:

[0256] Extraction module 1702, used to extract image information of each pixel in the target image;

[0257] A determination module 1704 is used to determine a pitch value corresponding to the image information of each pixel;

[0258] A conversion module 1706, for converting the image information of each pixel into melody track note data based on the pitch value;

[0259] The determination module 1704 is further used to determine the matching chord track note data based on the melody track note data or the music mode matching the target image;

[0260] The synthesis module 1708 is used to synthesize the melody track note data and the chord track note data to obtain an audio file.

[0261] The above-mentioned audio generation method, apparatus, computer equipment, storage medium and computer program product convert the image information of each pixel point into melody track note data based on the pitch value corresponding to the image information of each pixel point in the target image. The chord track note data is determined based on the music mode matching the melody track note data or the target image. The melody track note data and the chord track note data are synthesized into an audio file. The audio file is obtained by processing the image information of each pixel point in the target image. Compared with obtaining the audio file through a deep learning model trained by a large number of training samples, the time for collecting samples and model training can be saved, thereby improving the efficiency of generating audio.

[0262] In one embodiment, Fig.18 As shown, the notes are notes in a chord; the device also includes:

[0263] The determination module 1704 is further used to determine the music mode based on the image information of each pixel;

[0264] A first selection module 1710, configured to select a chord template from a chord template library according to a music mode;

[0265] The determination module 1704 is further used to determine the chord series corresponding to the image information of each pixel;

[0266] A first acquisition module 1712 is used to acquire the notes in the chord corresponding to the chord series in the chord template;

[0267] The determination module 1704 is further configured to determine a pitch value corresponding to the image information of each pixel point based on the notes in the chord.

[0268] In one embodiment, the apparatus further comprises:

[0269] The determination module 1704 is further used to determine the music mode based on the image information of each pixel;

[0270] The second acquisition module 1714 is used to acquire a target note set corresponding to the music mode;

[0271] The determination module 1704 is further configured to determine a pitch value corresponding to the image information of each pixel point based on each note in the target note set.

[0272] In one embodiment, the determination module 1704 is further configured to:

[0273] In the chord template corresponding to the music mode, the first chord constituent tone corresponding to the image information of each pixel is determined, and the chord track note data is generated based on the first chord constituent tone; or,

[0274] Based on the melody note data of each section in the melody track note data, the matching chord note data is determined to obtain the chord track note data composed of the chord note data; or,

[0275] A second chord constituent tone that is fixedly matched with the music mode is obtained, and chord track note data is generated based on the second chord constituent tone.

[0276] In one embodiment, the determination module 1704 is further configured to:

[0277] Select a chord template from the chord template library according to the music mode;

[0278] Determine the chord series corresponding to the image information of each pixel;

[0279] In the chord template, obtain the chord constituent notes corresponding to the chord series;

[0280] The chord constituent tone corresponding to the chord series is determined as the first chord constituent tone corresponding to the image information of each pixel point.

[0281] In one embodiment, the determination module 1704 is further configured to:

[0282] Determine the candidate chord corresponding to the melody note data of each section in the melody track note data;

[0283] Arrange and combine the alternative chords to obtain at least two alternative chord combinations;

[0284] In each alternative chord combination, the chord degree corresponding to the current alternative chord is used as a reference level, and adjacent alternative chords are scored until the scores corresponding to all alternative chords in each alternative chord combination are obtained;

[0285] Determine a combined score of each candidate chord combination based on the obtained scores;

[0286] Selecting a target chord combination from at least two candidate chord combinations based on the combination scores;

[0287] Chord note data matching the melody note data of each section is determined according to the target chord combination, and chord track note data composed of each chord note data is obtained.

[0288] In one embodiment, the determination module 1704 is further configured to:

[0289] Determine the weight value corresponding to each note data in the melody note data of each section;

[0290] Weighting each note data according to the weight value to obtain the note score of each note data;

[0291] Based on the note values, determining the note and value of the chord corresponding to each chord progression;

[0292] Sort the chords corresponding to each chord progression according to the note and value;

[0293] From among the chords, chords whose ranking reaches a preset ranking are selected as candidate chords.

[0294] In one embodiment, the apparatus further comprises:

[0295] A normalization module 1706 is used to perform normalization processing on the melody track note data to obtain normalized melody track note data;

[0296] The determination module 1704 is further used to weight each note data in the normalized melody note data according to the weight value to obtain the note score of each note data.

[0297] In one embodiment, the image information is a brightness value; the extraction module 1702 is further used to:

[0298] Get the chromaticity value of each pixel in the target image;

[0299] Determine the brightness value of each pixel in the target image based on the chromaticity value;

[0300] Determining the pitch value corresponding to the image information of each pixel includes:

[0301] Determine the pitch value corresponding to the brightness value of each pixel.

[0302] In one embodiment, the melody track note data includes at least two sections of melody note data, and each section of the melody note data includes at least two note data; the device further includes:

[0303] A merging module 1718 is used to merge the same note data that appear continuously in each section of the melody note data to obtain merged melody note data;

[0304] The determination module 1704 is further used to determine the note value of each note data in the merged melody note data of each section;

[0305] The conversion module 1706 is further used to convert the image information of each pixel into melody track note data based on the pitch value and the note value.

[0306] In one embodiment, the apparatus further comprises:

[0307] The third acquisition module 1720 is used to acquire media materials and generate a target video based on the media materials;

[0308] A second selection module 1722 is used to select a target video frame from the target video as a target image;

[0309] An adjustment module 1724, configured to adjust the aspect ratio of the target image according to the composition duration of the target video to obtain an adjusted target image;

[0310] The extraction module 1702 is further used to extract the image information of each pixel in the adjusted target image.

[0311] For the specific definition of the audio generating device, please refer to the definition of the audio generating method above, which will not be repeated here. Each module in the above audio generating device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0312] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.19As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store audio generation data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an audio generation method is implemented.

[0313] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig. 20 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an audio generation method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0314] Those skilled in the art will understand that Fig.19 , 20 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0315] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0316] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0317] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.

[0318] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0319] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0320] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. An audio generation method, characterized in that: The method comprises: Extract image information of each pixel in the target image; Determine a pitch value corresponding to the image information of each pixel; Converting the image information of each pixel into melody track note data based on the pitch value; Based on the melody track note data or the music mode matching the target image, the matching chord track note data is determined, including: in the chord template corresponding to the music mode, determining the first chord constituent tone corresponding to the image information of each pixel point, and generating the chord track note data based on the first chord constituent tone; or obtaining the second chord constituent tone fixedly matched with the music mode, and generating the chord track note data based on the second chord constituent tone; The melody track note data and the chord track note data are synthesized to obtain an audio file.

2. The method according to claim 1, characterized in that The note is a note in a chord; the method further comprises: Determining the music mode based on the image information of each of the pixels; Selecting a chord template from a chord template library according to the music mode; Determine the chord series corresponding to the image information of each pixel; In the chord template, obtaining the notes in the chord corresponding to the chord series; Determining the pitch value corresponding to the image information of each pixel point comprises: Based on the notes in the chord, a pitch value corresponding to the image information of each pixel point is determined.

3. The method according to claim 1, characterized in that The method further comprises: Determining the music mode based on the image information of each of the pixels; Obtaining a target note set corresponding to the music mode; Determining the pitch value corresponding to the image information of each pixel point comprises: Based on each note in the target note set, a pitch value corresponding to the image information of each pixel point is determined.

4. The method according to claim 1, characterized in that The step of determining matching chord track note data based on the melody track note data or the music mode matching the target image comprises: Based on the melody note data of each section in the melody track note data, the matching chord note data are determined to obtain the chord track note data composed of the chord note data.

5. The method according to claim 4, characterized in that Determining the first chord constituent tone corresponding to the image information of each pixel point in the chord template corresponding to the music mode includes: Selecting a chord template from a chord template library according to the music mode; Determine the chord series corresponding to the image information of each pixel; In the chord template, obtaining the chord constituent notes corresponding to the chord series; The chord constituent tone corresponding to the chord series is determined as the first chord constituent tone corresponding to the image information of each pixel point.

6. The method according to claim 1, characterized in that The step of determining matching chord note data based on the melody note data of each section in the melody track note data, and obtaining chord track note data composed of the chord note data includes: Determine the candidate chords corresponding to the melody note data of each section in the melody track note data; Arrange and combine the candidate chords to obtain at least two candidate chord combinations; In each of the alternative chord combinations, the chord degree corresponding to the current alternative chord is used as a reference level, and adjacent alternative chords are scored until the scores corresponding to all the alternative chords in each of the alternative chord combinations are obtained; Determine a combined score of each of the candidate chord combinations based on the obtained scores; Selecting a target chord combination from at least two candidate chord combinations based on the combination scores; Chord note data matching the melody note data of each section is determined according to the target chord combination, and chord track note data composed of the chord note data is obtained.

7. The method according to claim 6, characterized in that The step of determining the candidate chord corresponding to each section of the melody note data in the melody track note data comprises: Determine the weight value corresponding to each note data in the melody note data of each section; Weighting each of the note data according to the weight value to obtain a note score for each of the note data; Based on the note values, determining the note and value of the chord corresponding to each of the chord progressions; Sort the chords corresponding to the chord progressions according to the note sum values; From the chords, the chords whose ranking reaches a preset ranking are selected as the candidate chords.

8. The method according to claim 7, characterized in that The method further comprises: Normalizing the melody track note data to obtain normalized melody track note data; The step of weighting each note data according to the weight value to obtain the note score of each note data comprises: Each note data in the normalized melody note data is weighted according to the weight value to obtain the note score of each note data.

9. The method according to claim 1, characterized in that: The image information is a brightness value; the image information of each pixel in the target image is extracted including: Obtaining the chromaticity value of each pixel in the target image; Determine the brightness value of each pixel in the target image based on the chromaticity value; Determining the pitch value corresponding to the image information of each pixel point comprises: Determine the pitch value corresponding to the brightness value of each pixel point.

10. The method according to claim 1, characterized in that The melody track note data includes at least two sections of melody note data, and each section of the melody note data includes at least two note data; the method further includes: In the melody note data of each section, the same note data appearing continuously are merged to obtain merged melody note data; Determining the note value of each note data in the merged melody note data of each section; The converting the image information of each pixel point into melody track note data based on the pitch value comprises: Based on the pitch value and the note value, the image information of each pixel is converted into melody track note data.

11. The method according to any one of claims 1 to 10, characterized in that: The method further comprises: Acquire media material, and generate a target video based on the media material; Selecting a target video frame from the target video as the target image; Adjusting the aspect ratio of the target image according to the composition duration of the target video to obtain the adjusted target image; The step of extracting image information of each pixel in the target image comprises: The image information of each pixel is extracted from the adjusted target image.

12. An audio generating device, characterized in that: The device comprises: An extraction module, used to extract image information of each pixel in the target image; A determination module, used to determine the pitch value corresponding to the image information of each pixel point; A conversion module, used for converting the image information of each pixel into melody track note data based on the pitch value; The determination module is further used to determine the matching chord track note data based on the melody track note data or the music mode matching the target image, including: determining the first chord constituent sound corresponding to the image information of each pixel point in the chord template corresponding to the music mode, and generating the chord track note data based on the first chord constituent sound; or obtaining the second chord constituent sound fixedly matched with the music mode, and generating the chord track note data based on the second chord constituent sound; The synthesis module is used to synthesize the melody track note data and the chord track note data to obtain an audio file.

13. The device according to claim 12, characterized in that The notes are notes in a chord; the device also includes: The determination module is further used to determine the music mode based on the image information of each pixel point; A first selection module, used for selecting a chord template from a chord template library according to the music mode; The determination module is further used to determine the chord series corresponding to the image information of each pixel point; A first acquisition module, used for acquiring the notes in the chord corresponding to the chord series in the chord template; The determination module is further used to determine the pitch value corresponding to the image information of each pixel point based on the notes in the chord.

14. The device according to claim 12, characterized in that The device also includes: The determination module is further used to determine the music mode based on the image information of each pixel point; A second acquisition module is used to acquire a target note set corresponding to the music mode; The determination module is further used to determine the pitch value corresponding to the image information of each pixel point based on each note in the target note set.

15. The device according to claim 12, characterized in that The determination module is also used to determine matching chord note data based on the melody note data of each section in the melody track note data, and obtain chord track note data composed of the chord note data.

16. The device according to claim 15, characterized in that The determination module is also used to select a chord template from a chord template library according to the music mode; determine the chord series corresponding to the image information of each pixel point; obtain the chord constituent tones corresponding to the chord series in the chord template; and determine the chord constituent tones corresponding to the chord series as the first chord constituent tones corresponding to the image information of each pixel point.

17. The device according to claim 12, characterized in that The determination module is also used to determine the alternative chords corresponding to the melody note data of each section in the melody track note data; arrange and combine the alternative chords to obtain at least two alternative chord combinations; in each of the alternative chord combinations, use the chord degree corresponding to the current alternative chord as a reference level to score the adjacent alternative chords until the scores corresponding to all the alternative chords in each of the alternative chord combinations are obtained; determine the combination score of each of the alternative chord combinations based on the obtained scores; select a target chord combination from at least two of the alternative chord combinations based on the combination scores; determine the chord note data that matches the melody note data of each section according to the target chord combination, and obtain the chord track note data composed of the chord note data.

18. The device according to claim 17, characterized in that The determination module is also used to determine the weight values ​​corresponding to each note data in the melody note data of each section; weight each note data according to the weight value to obtain the note score of each note data; determine the note sum value of the chord corresponding to each chord series based on the note score; sort the chords corresponding to each chord series according to the note sum value; and select the chord whose sorting ranking reaches a preset ranking from each chord as the alternative chord.

19. The device according to claim 18, characterized in that The device also includes: A normalization module, used for normalizing the melody track note data to obtain normalized melody track note data; The determination module is further used to weight each note data in the normalized melody note data according to the weight value to obtain the note score of each note data.

20. The device according to claim 12, characterized in that The image information is a brightness value; The extraction module is further used to obtain the chromaticity value of each pixel in the target image; Determine the brightness value of each pixel in the target image based on the chromaticity value; The determining of the pitch value corresponding to the image information of each pixel point includes: determining the pitch value corresponding to the brightness value of each pixel point.

21. The device according to claim 12, characterized in that The melody track note data includes at least two sections of melody note data, and each section of the melody note data includes at least two note data; the device also includes: A merging module, used for merging the same note data that appear continuously in the melody note data of each section to obtain merged melody note data; The determination module is further used to determine the note value of each note data in the merged melody note data of each section; The conversion module is further used to convert the image information of each pixel into melody track note data based on the pitch value and the note value.

22. The device according to any one of claims 12 to 21, characterized in that The device also includes: A third acquisition module, used to acquire media materials and generate a target video based on the media materials; A second selection module is used to select a target video frame from the target video as the target image; An adjustment module, used for adjusting the aspect ratio of the target image according to the composition duration of the target video to obtain the adjusted target image; The extraction module is further used to extract image information of each pixel in the adjusted target image.

23. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.

24. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

25. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Automatic harmony method, device, and terminal automatic harmony operation method

    CN105161087A

  • Music style conversion method, music style conversion device and terminal equipment

    CN110246472A

  • Music generation method and device

    CN110444185A